Grok has the best value for coding
Grok has the best value for codingCognition: The FrontierCode leaderboard is now live: a dedicated page that tracks which models are writing code you’d actually merge.…
Grok has the best value for codingCognition: The FrontierCode leaderboard is now live: a dedicated page that tracks which models are writing code you’d actually merge.…
in line with my impression so far5.5, not 5.6But holy shit, 5.5 is not that oldArtificial Analysis: Kimi K3 in Kimi Code CLI scores 57 and…
RT cheatywhat am i even looking at? is this a joke? how was this released two months ago and you're only posting abotu it now? the…
Underrated takeYes, "naturally" Chinese labs should stick to models on the level of DSV4-Flash. Super cheap, decent-ish. Economy class. Spend the surplus on marketing and "Application…
The funny thing is if K3 is heavily distilled from *Opus* it dunks onAlex Kolicich: It looks likely that Kimi K3 was heavily distilled from FableIf…
btw if you havent set your {codex | claude | gemini | devin} automations to autoresearch how to improve your seo/aeo every week you are really…
this is cool:https://welcome-to-codex.openai.chatgpt.site/
Re Here's how to add the Codex Security plugin in Codex and get started:Add the plugin in Codex. After installation is complete, the button changes to…
GPT-5.6 Sol sets a new state of the art in cybersecurity on “The Last Ones” cyber range.We’re already seeing that capability translate into defensive outcomes: helping…
RT Min ChoiGrok 4.5 is only 1.5T parameters.Kimi K3 is 2.8T... and costs 3× more per task.SpaceXAI cracked intelligence efficiency.Now imagine Grok at 3T.Elon Musk: Grok…
it's funny how @TMTLongShort's bio starts with "Economic Blitzkrieg". Unfortunate naming. I wonder if Nazis also thought they're losing because they're "pussyfooting around", being too Merciful…
RT DeepBurnerAdded kimi-3 to my Surface Evolver Bench. It scores on par with Fable 5 and Sol (xhigh). The cost is in fact higher than Sol,…
RT Laude InstituteAndy @andykonwinski and @swyx talked through why AI researchers may be uniquely well-positioned to wield power responsibly right now, why most of them are…
RT Rothmus 🏴On this day in 1918 the Bolsheviks herded Tsar Nicholas II, Empress Alexandra, and their five children into a basement in Yekaterinburg and opened…
The bagholder mindset is fearsome to behold> We have already boarded a train we cannot get off—one that leaves us no choice but to keep pouring…
RT AfterQueryKimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5. An open weight model now outperforms all closed-sourced models. Read more in the…
Sol gets the thing done:tobi lutke: OpenAI shipped /goal not too long ago. I feel like GPT-5.6-sol is the first model that just doesn't need it…
Still building your #SIGGRAPH2026 agenda? We've got you covered. 🤝👉 Learn the latest research breakthroughs at the NVIDIA keynote.👉 Build new skills in hands-on labs.👉 Explore…
China is really obsessed with material sciencethey graduate buttloads of these researchers, publish like 60% of the world's total high impact material science and engineering researchand…
GPT-5.6 Sol is the state of the art in cyber. Seeing significant results in applying it to finding and fixing novel vulnerabilities.Sign up as a defender…
The issue isn't so much about logits, this is a red herring imothe main problem is on-policy vs off-policy RL. Chinese labs are shifting towards MOPD.…
10,000 reasons people love GPT-5.6 Sol https://switch-to-codex.openai.chatgpt.site/Tibo: Or… what if we gave you $100 in Codex credits if you tell us what you love about GPT-5.6…
I really want to see more info on this new Ascend configuration@zephyr_z9 do you hear anything from WAIC?Kyle Chan: China’s message at this year’s World AI…
it's remarkable how people don't reflect on the way differential credulity shapes their views.Some shitty handwavy Reuters report: a credible source on "real" Chinese position on…