Prefer Grok's outputs to GPT 5.5 by some margin
Prefer Grok's outputs to GPT 5.5 by some margin
Every primary-source story across every tracked model. Filter by clicking a chip.
Prefer Grok's outputs to GPT 5.5 by some margin
Show HN: Runtime authorization for Claude Code, Cursor, and CodexHi HN, Fernando and I built Kastra. Kastra intercepts AI agent tool calls and evaluates them against…
"OpenAI has arrived at a new solution that seems to combine the best parts of the two solutions that human top performers ultimately reached."Yoichi Iwata: AWTF…
Claude’s new Reflect dashboard doesn’t just visualize how you use AI. It also subtly reinforces how much of your daily work now depends on Anthropic’s chatbot.
Three big AI IPOs are set to generate more value than all the U.S. VC backed exits since 2000.
My only worry is that DeepSeek itself has NIH syndrome after so much of their work became foundational to this eraThey adopt third party tricks, but…
RT Tyler Brunogrok 4.5 is seriously good. it's going in the direction which i'm hoping to see more of -- an emphasize on speed as well…
RT Alexandr Wang1/ muse spark 1.1 is an industry-competitive agentic and coding model. across many agentic evals it rivals gpt-5.5 and opus-4.8.available now through the new…
Introducing Muse Spark 1.1 and the Meta Model API
Your Prompts and Skills need a system of record.
> almost consistently outperformsGRPO is a strong, extensible baseline and I suspect that people dunking on it (and DeepSeek) are motivated by personal mathematical aesthetics and…
The popularity of Spotify Wrapped has kicked off a wide range of year-in-review features, on apps from YouTube to Uber - and now, the lookback trend…
OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent of its tasks are broken. The company…
More Fable, GPT-5.6 and new ChatGPT voice
Learn how GPT-5.6 powers Microsoft 365 Copilot with stronger AI capabilities across Word, Excel, PowerPoint, Chat, and Cowork for faster, higher-quality work.
RT GrokTry Grok 4.5 for free, an all new Opus-class model that is fast and low cost. Great for real-world coding and engineering tasks.
RT GrokTry Grok 4.5 for free, an all new Opus-class model that is fast and low cost. Great for real-world coding and engineering tasks.
RT GrokGrok 4.5 is built for real-world engineering. It excels in large codebases and handles long-running tasks that span multiple repositories, hundreds of skills, and a…
I made this simple 3D Geometry Wars-style game using my coding agent, Jarvis Code, with GLM 5.2. You can play it here: https://jarvis-llm-codec.github.io/jarvis-code/geometry-wars-3d.html I was honestly…
submitted by /u/beneath_steel_sky [link] [comments]
Databricks benchmarked coding agents on its own multi-million-line codebase and found that the Chinese open-source model GLM 5.2 matched Anthropic's Opus 4.8 at $1.28 per task…
RT gumto put some context on what imo ppl should infer from this result. i would have expected a similarly scaled up specialized model (like OpenAi's…
Details about the OpenAI Bio Bounty program
ChatGPT Work is an agent that can take action across your apps and files, stay with a project for hours if needed, and turn a goal…
We track 28 AI models across text, image, video, and audio domains. Each model has its own filtered news feed — click a chip above to see only that model's primary-source coverage. Tagging is automatic at ingest time using a strict title-only keyword match (so a benchmark post that mentions five models in its summary only shows up under whichever model is named in the headline — no thin-content drift).
Text / LLMs: GPT, Claude, Gemini, Gemma, Llama, Mistral, Grok, Qwen, DeepSeek, Kimi, GLM, MiniMax, Yi, Hunyuan, Command, Phi.
Image generation: FLUX, Stable Diffusion, Midjourney, Imagen.
Video generation: Sora, Veo, Runway, Luma, Kling, Pika.
Audio / voice: ElevenLabs, Suno.