Unbelievable MiniMax and bone W
Unbelievable MiniMax and bone Wbone: Hakuna Matata starring Nikita Boar, Elon Tusk, and Grok
Every primary-source story across every tracked model. Filter by clicking a chip.
Unbelievable MiniMax and bone Wbone: Hakuna Matata starring Nikita Boar, Elon Tusk, and Grok
I previously demonstrated an anthropic impossibility theorem, showing that in Duplicates Sleeping Beauty, there was no possible probability theory that obeyed both the martingale condition and…
xAI has released Imagine Image 2.0 as a new image generator for Grok. The model ranks second in the Arena benchmarks, just behind OpenAI's GPT-Image-2. New…
I've been tuning my new Radeon AI Pro R9700, and figured that this would be useful information for people who are trying to optimise their setups.…
$10,000 kill my saas in a weekend competition is live!TECH STACK:any coding agentany modelup to $500 in token spend incl subscriptionssee luma for description. join waitlist…
I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real. Everyday tasks with Hermes agent? Effortless. Coding…
I’ve been using 3.6 27B Q4, and that quant is fast on an M5. The code has been average, but consistently “good enough.” And, after a…
I compared Qwen 35B-A3B MoE against Qwen 27B dense on a series of local coding-maintenance tasks. On my R9700/llama.cpp setup, the MoE model generated about 3.9×…
You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff.If…
Mistral AI has released Shieldstral 1.0 3B, an open-weights, policy-adaptive multimodal safety classifier that frames content moderation as a single yes/no question instead of a fixed…
RT GrokAnnouncing Imagine Image 2.0, our next generation image model with precision editing, crisp text rendering, improved factuality, and real world usefulness.Image 2.0 helps you make…
chatgpt work for removing the toiljason: its one of those days with ChatGPT Work
Can’t think of a single thing that new Flash does worse than V4-Flash-Preview. It’s pretty rare actually. Like Sonnet 3 to 3.5. It’s not more verbose…
ChatGPT team is shipping:Adam Fry: This week's ChatGPT feature drop - Aug 7:1/ Rich formatting in our web composer – When you paste in emails or…
I didn’t expect thisMiniMax is absolutely much stronger in videogen than in LLMs right now. To think they can even open source such a modelChetaslua: Holy…
We analyzed DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE.A DeepSeek-first cascade with test-suite verification solved MORE tasks than Luna alone at 37% lower cost per…
Re @OpenAI oo claude code has this now!!! need to tryhttps://x.com/ClaudeDevs/status/2085817074816070014ClaudeDevs: New in Claude Code: your sessions can now message each other.Instead of having to re-explain…
dear openaijust make a new phoneeveryone wants openaiphonewe can read 2-4x faster than we talk and speakopenai alexa reachy hybrid is fine but pls just be…
RT Dandid you know Grok is the only AI you can also use as a verb
OpenAI gave a last-minute presentation at the Black Hat security on Wednesday about "the Hugging Face Incident" (previously on this blog). The video was published yesterday.…
Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory)…
RT ℏεsam🚨 BREAKING — Anthropic investors worry Dario Amodei’s AI doom marketing could hurt its upcoming IPO.“He’s more of a religious leader than he is a…
OpenAI said it has suspended work on some aspects of its upcoming model Astra over concerns about its cybersecurity prowess.
Speaking of, I wonder if DeepSeek will fail the FelonyBench. Unlike OpenAI, they pay a lot of attention to using up 100% of hardware potential. Less…
We track 28 AI models across text, image, video, and audio domains. Each model has its own filtered news feed — click a chip above to see only that model's primary-source coverage. Tagging is automatic at ingest time using a strict title-only keyword match (so a benchmark post that mentions five models in its summary only shows up under whichever model is named in the headline — no thin-content drift).
Text / LLMs: GPT, Claude, Gemini, Gemma, Llama, Mistral, Grok, Qwen, DeepSeek, Kimi, GLM, MiniMax, Yi, Hunyuan, Command, Phi.
Image generation: FLUX, Stable Diffusion, Midjourney, Imagen.
Video generation: Sora, Veo, Runway, Luma, Kling, Pika.
Audio / voice: ElevenLabs, Suno.