RT Psyho: All problems have been solved by OpenAI!
RT PsyhoAll problems have been solved by OpenAI!
Every primary-source story across every tracked model. Filter by clicking a chip.
RT PsyhoAll problems have been solved by OpenAI!
Try out Grok 4.5!Boris Skorobogaty: Grok 4.5 is finally the model that can handle the entire workflow end-to-end:/design /execute-plan --effort --instructions "your specific guidance"/pr-babysit add /loop…
SDK regeneration (#808) * [fern-generated] Update SDK Generated by Fern CLI Version: unknown Generators: - fernapi/fern-python-sdk: 4.42.0 * [fern-replay] Applied customizations Patches with unresolved conflicts (1):…
RT SimilarwebTraffic to http://grok.com more than doubled year-over-year in H1.
Grok 4.5 is also rank 1 in SWE marathonTech Dev Notes: SpaceXAI has added SWE Marathon benchmark to Grok 4.5 blog Grok 4.5 ranks #1 in…
Looks like Grok 4.5 is #1 on at least a few benchmarks. Better than expected.Thibault Jaigu: Very impressed by grok-4.5! Awesome job @elonmusk & @mntruell Best…
RT Arthur MacWatersImpressive Intelligence vs. Cost for Grok 4.5I think given terafab + the Tesla FSD team's expertise in making very efficient AI models, this advantage…
Grok is making progressleo 🐾: lmao wtf
Cool that Grok 4.5 is #1 in some respects, even with respect to Fable 5Artificial Analysis: SpaceXAI's Grok 4.5 takes the #1 spot on AutomationBench-AA with…
I don't know where this is headed, but I don't like it. https://futurism.com/artificial-intelligence/open-source-ai-model-scary-mythos GLM-5.2 can be downloaded by anybody, can be run on virtually any hardware,…
SpaceXAI continues to move faster than any other frontier lab on earth.
**OpenAI** launched the **GPT-5.6** family with three models: **Sol**, **Terra**, and **Luna**, integrated across **ChatGPT**, **Codex**, and the API. Pricing tiers range from **$1 to $5…
RT Yun-Ta TsaiI appreciate the many xAI and Cursor engineers who dedicated their time to addressing feedback from Tesla.I remember meeting Andrew a few weeks after…
Since I am needling every model maker tonight about minor but important issues, one more: Grok 4.5 has no model card. Companies that are trying to…
Note: the modeling assumptions and conclusion are Thomas Kwa's opinion, and others at METR disagree. [1] Also, the math was checked by Claude but not a…
The metrics discussion at OpenAI is a little confusing to me. I appreciate the clarification about bad benchmarks, but they spent a lot of money developing…
arXiv:2501.10870v2 Announce Type: replace-cross Abstract: The principal objective of this work is twofold within nonparametric regression settings: (1) to establish the minimax optimal convergence rates for…
Hey Everyone, here is an update on MTPLX! One month after releasing MTPLX V1 which brought a swift based app and upgraded CLI for coding use…
RT GrokUse Grok 4.5 to build full-stack apps with ConvexMikeysee: Wow Grok 4.5 is very impressive at @convex code. Almost perfect score with a very very…
RT Artificial AnalysisSpaceXAI's Grok 4.5 takes the #1 spot on AutomationBench-AA with a score of 51%, ahead of Claude Fable 5 (49%) and Claude Opus 4.8…
RT Kun Chengrok 4.5 made me give grok build a serious run todayhere's my honest first impression (non affiliated neutral view point):1. grok build is a…
RT grimGrok Build mogs Claude Code & Codex TUIs i can't lie. Not to mention the perf of Grok 4.5...It's not even a competition at this…
GPT-Live is now fully rolled out to all ChatGPT users on Go, Plus, and Pro plans. Free user rollout is in progress. Update to the latest…
Grok speaks like a robot. Actually like the @grok doing responses here. Terse, stilted, enumerating options. ESL vibes. Agent GrokI like it. totally useless for many…
We track 28 AI models across text, image, video, and audio domains. Each model has its own filtered news feed — click a chip above to see only that model's primary-source coverage. Tagging is automatic at ingest time using a strict title-only keyword match (so a benchmark post that mentions five models in its summary only shows up under whichever model is named in the headline — no thin-content drift).
Text / LLMs: GPT, Claude, Gemini, Gemma, Llama, Mistral, Grok, Qwen, DeepSeek, Kimi, GLM, MiniMax, Yi, Hunyuan, Command, Phi.
Image generation: FLUX, Stable Diffusion, Midjourney, Imagen.
Video generation: Sora, Veo, Runway, Luma, Kling, Pika.
Audio / voice: ElevenLabs, Suno.