Effective harnesses for long-running agents
Agents still face challenges working across many context windows. We looked to human engineers for inspiration in creating a more effective harness for long-running agents.
Every primary-source story across every tracked model. Filter by clicking a chip.
Agents still face challenges working across many context windows. We looked to human engineers for inspiration in creating a more effective harness for long-running agents.
add evaluation reproduction code for MMMU,MathVision,ODinW-13,RealWorldQA
FLUX.2 brings professional-grade image generation and editing with unprecedented detail, multi-reference support, and enterprise efficiency.
The most capable model in Windsurf yet, now available at Sonnet prices for a limited time
We’ve added three new beta features that let Claude discover, learn, and execute tools dynamically. Here’s how they work.
Start-ups like Conservation X Labs are using Meta’s Segment Anything Models to bolster on-the-ground conservation expertise.
Merge pull request #301 from TianQi-777/patch-5 Update README_zh.md
Merge pull request #300 from TianQi-777/patch-4 Update README.md
ExecuTorch, Meta’s open source, lightweight, and efficient inference engine, has been instrumental in enabling on-device AI capabilities across Meta’s family of apps.
Bringing the next generation of tool-calling agents to the xAI API
Announcing Our Landmark Partnership with Saudi Arabia and HUMAIN
This release introduces two new state-of-the-art models: SAM 3D Objects for object and scene reconstruction, and SAM 3D Body for human body and shape estimation.
From chatbots to agents
fix patch size for video in image list.
Grok 4.1 is now available to all users on grok.com, 𝕏, and the iOS and Android apps. It is rolling out immediately in Auto mode and…
We are pleased to offer a HIPAA-compliant Business Associate agreement (BAA) to enable customers across the healthcare industry to work with us to develop secure custom…
GPT 5.1, GPT 5.1-Codex, and GPT-5.1-Codex Mini deliver a solid upgrade for agentic coding with variable thinking and improved steerability
Merge pull request #42 from rogeryoungh/patch-2 Reformat license
Merge pull request #41 from rogeryoungh/patch-1 Update LICENSE
Introducing Shared Memory IPC Caching — a high-performance caching mechanism contributed by Cohere to the vLLM project.
Merge pull request #1757 from ZLkanyo009/main [FIX] fix README sglang server part
We’re introducing Meta Omnilingual Automatic Speech Recognition, a suite of models providing automatic speech recognition capabilities for over 1,600 languages.
Merge pull request #32 from MiniMax-AI/feat/issue-template update issue template
Merge pull request #24 from MiniMax-AI/feat/issue-template
We track 28 AI models across text, image, video, and audio domains. Each model has its own filtered news feed — click a chip above to see only that model's primary-source coverage. Tagging is automatic at ingest time using a strict title-only keyword match (so a benchmark post that mentions five models in its summary only shows up under whichever model is named in the headline — no thin-content drift).
Text / LLMs: GPT, Claude, Gemini, Gemma, Llama, Mistral, Grok, Qwen, DeepSeek, Kimi, GLM, MiniMax, Yi, Hunyuan, Command, Phi.
Image generation: FLUX, Stable Diffusion, Midjourney, Imagen.
Video generation: Sora, Veo, Runway, Luma, Kling, Pika.
Audio / voice: ElevenLabs, Suno.