Qwen 3.8 Max now ranked as best overall model ahead of Opus 5 by Artificial Analysis agentic index
submitted by /u/anderspitman [link] [comments]
Every primary-source story across every tracked model. Filter by clicking a chip.
submitted by /u/anderspitman [link] [comments]
Article URL: https://artificialanalysis.ai/?intelligence=agentic-index Comments URL: https://news.ycombinator.com/item?id=49200652 Points: 144 # Comments: 60
OpenAI has updated GPT-5.6 Sol with more focused responses and a reasoning slider that lets users adjust how deeply the model thinks. Free users will get…
Re Plus and Pro users can access the updated version of GPT‑5.6 Sol in ChatGPT along with the new slider starting today.This version of GPT-5.6 Sol…
Re Plus and Pro users also now have a slider to choose how much reasoning effort ChatGPT puts into each response.We think it’s easier to use,…
Re In addition to the upgrade in intelligence with GPT-5.6 Luna, Free and Go users can now use the “Think” button for more reasoning on harder…
We’re making better intelligence easier to access in ChatGPT for everyone:- GPT-5.6 Sol now powers both Instant and deep reasoning for Plus & Pro users, delivering…
Re The new GPT-5.6 Sol powers all chats for paid users, including Instant, creating one consistent experience. In our high-stakes factuality evaluation covering finance, medicine and…
Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users
How the world is putting ChatGPT to work
RT Visual Studio Code📣 Kimi K3 is now available in GitHub Copilot for @code!Try Kimi Moonshot's latest open-weight model for agentic coding, now hosted by @FireworksAI_HQ.📖…
Suno announced plans to implement a new watermarking technology and download policy to limit the spread of spammy AI tracks and increase transparency. In a lengthy…
Microsoft generated $24.1 billion in AI revenue through OpenAI in the fiscal year ending in June. That's about 70 percent of its total AI business, according…
OpenAI said that ChatGPT free and Go users are also getting a new think button for complex queries.
RT GitHub📣 @Kimi_Moonshot's Kimi K3, an open-weight model, is now generally available and rolling out in GitHub Copilot. The model shows frontier-level abilities on agentic coding…
Kimi K3, an open-weight model, is now generally available in GitHub Copilot. The model shows frontier-level abilities on agentic coding with highly cost-effective pricing. Kimi K3…
Link to the article: KV Cache Quantization Benchmarks: KVarN, Precision Tail KLD benchmarks with BeeLlama.cpp v0.4.0, fork of llama.cpp with more KV cache quantization options. Models:…
OpenAI is making a big change for ChatGPT users on its free and Go tiers: starting next week, users on those tiers will be able to…
(Disclaimer: I am a noob and don’t know what I am doing) Gemma 4 31b unsloth/gemma-4-31B-it-qat-GGUF I took the f16 MTP draft model and quantised it…
Composio tested Deepseek V4 Flash across four agent frameworks on 30 real-world tasks. Success rates were mostly similar, but costs varied by nearly 3x: OpenCode came…
Kimi K3 scores nearly TWICE Claude Fable 5 on Harvey LAB-AA's hard autonomous legal tasks
RT Amir HaghighatWe're the fastest provider for DeepSeek V4 Flash and Kimi K3 on @huggingface, which measures based on real production traffic vs test data.Baseten: Baseten…
A regulated customer needed all Claude Code inference processed in a single AWS Region (London), not just in-geography. This post shows two ways to pin Claude…
Following up on yesterday's post about running everyone's faves on 2 x 16gb cards while maximizing performance and KV. Previous post data used abandoned Cu130 VLLM…
We track 28 AI models across text, image, video, and audio domains. Each model has its own filtered news feed — click a chip above to see only that model's primary-source coverage. Tagging is automatic at ingest time using a strict title-only keyword match (so a benchmark post that mentions five models in its summary only shows up under whichever model is named in the headline — no thin-content drift).
Text / LLMs: GPT, Claude, Gemini, Gemma, Llama, Mistral, Grok, Qwen, DeepSeek, Kimi, GLM, MiniMax, Yi, Hunyuan, Command, Phi.
Image generation: FLUX, Stable Diffusion, Midjourney, Imagen.
Video generation: Sora, Veo, Runway, Luma, Kling, Pika.
Audio / voice: ElevenLabs, Suno.