Grok
GrokTech Dev Notes: Grok 4.5 is #2 in FrontierSWE benchmark
Every primary-source story across every tracked model. Filter by clicking a chip.
GrokTech Dev Notes: Grok 4.5 is #2 in FrontierSWE benchmark
Good reason to try Grok 4.5 with Grok Build. It gets better every day!X Freeze: Grok 4.5 just took the #1 spot on the Long-Horizon Terminal-Bench,…
same energythere's no hard qualitative boundary between "Erdos problem" and "Putnam problem"OpenAI just invests into doing these publicity stuntsI would obviously love it if DeepSeek claimed…
Can a handful of engineers really do the work of an army of consultants? That’s the bet behind Ode with Anthropic — the joint venture dedicated…
> Mind you, DeepSeek-V4, GLM-5.2 or Kimi-K2.6/2.7 haven't solved a single Erdös problem so far.Ok this is enough Lisan. Mind you, not all Erdös problems are…
One of great problems for DeepSeek is that great priorities for research and great product are not aligned. They can afford ad hoc harnesses, slow (but…
Article URL: https://dpa-international.com/economics/urn:newsml:dpa.com:20090101:260715-930-389143/ Comments URL: https://news.ycombinator.com/item?id=48921461 Points: 37 # Comments: 13
How I tricked Claude into leaking your deepest, darkest secrets I've been impressed by the way the Claude web_fetch tool is designed to avoid data exfiltration…
another DeepGEMM updatewhat is funny is that DeepSeek's efficiency (and margin) is a moving target. GLM can optimize DSA, why can't they keep pushing their V4…
RT Tesla Owners Silicon ValleyBREAKING: Grok 4.5 has climbed to #2 on the FrontierSWE benchmark.The result places Grok 4.5 among the world's top-performing AI models for…
I set out to make Parakeet speech-to-text faster on CPU. Not a GPU demo. Just a real baseline and a keep gate. I did not want…
hey HN - Claude pre-created users in Clerk with null emails/names as "guest users" on a contract job. Wasn't in any plan. The CTO asked why,…
The creators behind the magic are here. ✨From wild ideas to award-winning creations, discover the stories behind the works that stood out at Kling AI NEXTGEN…
> DeepSeek's gross margin from selling v4 API access is...70% to 80% (!)I wonder if Wenfeng would still say «not to subsidize, nor to reap excessive…
Anthropic-backed Ode launches as AI labs bet that embedding forward-deployed engineers inside enterprises is the key to accelerating enterprise AI adoption.
Don't do this to me pleaseI don't want to feel disappointed in the WhaleGLM 5.2 level is enough to do well at these costsStretch goal: 5.6…
RT X FreezeGrok just leveled up againTwo major new connectors have been added:• Stripe → payments, customers and invoices• Calendly → availability, scheduling and meetingsGrok can…
> DeepSeek's gross margin from selling v4 API access is...70% to 80% (!)I wonder if Wenfeng would still say «not to subsidize, nor to reap excessive…
RT Peter H. Diamandis, MDApple just sued OpenAI for trade secret theft in a 41 page federal complaint. The two companies were partners a year and…
OpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI.
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
Since early June, OpenAI's coding tool Codex encrypts the instructions a main agent passes to its subagents. Developers can no longer track how tasks get delegated…
Frontier models are just so good though. Fable 5... Gemini 3.1 Pro for design critique and brainstorming. Grok for verification passes. Antigravity with Gemini 3.5 Flash…
OpenAI plans to enter hardware with a portable, screenless smart speaker. Equipped with a camera, sensors, and moving mechanical parts, the device is designed as an…
We track 28 AI models across text, image, video, and audio domains. Each model has its own filtered news feed — click a chip above to see only that model's primary-source coverage. Tagging is automatic at ingest time using a strict title-only keyword match (so a benchmark post that mentions five models in its summary only shows up under whichever model is named in the headline — no thin-content drift).
Text / LLMs: GPT, Claude, Gemini, Gemma, Llama, Mistral, Grok, Qwen, DeepSeek, Kimi, GLM, MiniMax, Yi, Hunyuan, Command, Phi.
Image generation: FLUX, Stable Diffusion, Midjourney, Imagen.
Video generation: Sora, Veo, Runway, Luma, Kling, Pika.
Audio / voice: ElevenLabs, Suno.