This is a big jump in ARC-AGI-3.
This is a big jump in ARC-AGI-3.ARC Prize: Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2%The previous high score (7.8%) was set…
This is a big jump in ARC-AGI-3.ARC Prize: Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2%The previous high score (7.8%) was set…
Opus 5 replaced Opus 4.8 for me, it was generally stronger in everything... except it shares some of the weird language quirks of Fable, including a…
Unexpected finding from a big study on what ChatGPT did to colleges: "once the COVID-19 disruption is modeled separately, the introduction of ChatGPT had no detectable…
As a joke I prompted Codex "Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is…
Glad to see Google sharing data on how Gemini is being used. Especially interesting is that the usefulness of multimodal AI for manual labor may be…
I wrote the latest of my occasional guides to which AI to use right now for non-experts who want to get stuff done.The agentic systems available…
Wish I had a list of unsolved math problems handy. (But the sudden flow of math proofs/disproofs is a good indication of what will happen in…
“Generate a fake, but believable, witty Churchill insult at a party and explain the context. It should be very clever and original”This time, I think GPT…
Who would you give authorship to? The person who wrote 58 words of prompts, or GPT-5.6 Pro?Dmitry Rybin: Dinitz-Garg-Goemans conjecture is false. This graph theory problem…
I have no inside information, but every sign so far is that there is growing tension about open weights models between the US & China: the…