Grok 4.5 even ranks slightly above Fable, which is an incredibly good model, on some software benchmarks!
Grok 4.5 even ranks slightly above Fable, which is an incredibly good model, on some software benchmarks!
Every primary-source story across every tracked model. Filter by clicking a chip.
Grok 4.5 even ranks slightly above Fable, which is an incredibly good model, on some software benchmarks!
submitted by /u/fallingdowndizzyvr [link] [comments]
Grok 4.5 is Opus class for browser useAlexander Yue: I take it all back. We just got access to eval Grok 4.5 and it has landed…
A man reading, minding his own business. Then things start to get a little out of proportion. The Luma Skill behind it, by Eli Coleman. Made…
i'd love to see interesting things people have built with 5.6 sol.i will send the person who made the coolest thing a special gift from the…
RT elvisI don't think Anthropic realizes how disruptive these changes are to users. I appreciate the extension, but please stop playing games. Either keep it under…
RT Miki OvshievichWell shit - grok 4.5 is KILLING it it in my video studio evals!6/33 -> 23/33 with ridiculous cost efficiency compared to 5.6 models
This started based off of a hunch. We usually use OpenCode, but were 'forced' to use Claude Code for a while due to issues with Meridian.…
Article URL: https://ploy.ai/blog/migrating-a-production-ai-agent-to-gpt-5-6 Comments URL: https://news.ycombinator.com/item?id=48882716 Points: 6 # Comments: 1
RT @gailalfaratx: Grok 4.5 is the truth seeking leader that other AIs can aspire to become
It's annoying that you can't paste a link to a (shared) Claude transcript into a Claude Code session, because Anthropic's anti-scraping measure prevent its own tools…
Claude Code now has a built-in browser that lets the AI open, read, and interact with web pages directly inside the development environment. Write actions on…
Article URL: https://www.elliotcsmith.com/autoresearch-claude-and-constrained-optimization/ Comments URL: https://news.ycombinator.com/item?id=48881498 Points: 3 # Comments: 0
RT PJ AceEveryone said AI ads would never win at Cannes.Two of them just did. Both built on Kling.Here's why it's a bigger deal than the…
Anthropic’s research on Claude found a silent internal workspace they call J-space — hidden reasoning that never shows up as visible text. Classic example: the model…
As many of you know t/s is super important. It's how fast your stuff gets done. I create via open code benchtest and run it. Thanks…
Let's start a discussion about what can be done to make local models more reliable. I've been using Qwen3.6-27B a lot lately, and have noticed the…
S&P Global has downgraded Oracle's credit rating to "BBB-," one notch above junk status. OpenAI accounts for roughly half of Oracle's $638 billion in contractual obligations.…
I recently posted some posts with VLLM showing issues with TTFT and concurrency with 4x 5060 ti's. Wanted to share this benchmark to provide what worked…
The Wall Street Journal printed an outright false headline and heavily misleading story claiming this, which of course was uncritically amplified by the usual suspects. I…
Try Grok 4.5 and see for yourselfMin Choi: Ok Grok 4.5 is insane.People are already building things that shouldn't be possible this fast.10 wild examples:
I’ve been playing around with Opencode and realized how 70% of the capability of my model comes from the agents I can use rather than the…
Anthropic analyzed 1.2 million Claude Cowork sessions from more than 600,000 organizations. About half of all usage goes toward business processes and text creation, what Anthropic…
OpenAI CEO Sam Altman now says he's "pretty sure" AI has created more jobs than it's eliminated. That's a sharp turn from his earlier warnings about…
We track 28 AI models across text, image, video, and audio domains. Each model has its own filtered news feed — click a chip above to see only that model's primary-source coverage. Tagging is automatic at ingest time using a strict title-only keyword match (so a benchmark post that mentions five models in its summary only shows up under whichever model is named in the headline — no thin-content drift).
Text / LLMs: GPT, Claude, Gemini, Gemma, Llama, Mistral, Grok, Qwen, DeepSeek, Kimi, GLM, MiniMax, Yi, Hunyuan, Command, Phi.
Image generation: FLUX, Stable Diffusion, Midjourney, Imagen.
Video generation: Sora, Veo, Runway, Luma, Kling, Pika.
Audio / voice: ElevenLabs, Suno.