AI Daily Brief — 24 November 2025
Monday landed the biggest model launch of November. Anthropic released Claude Opus 4.5 — first model to break 80% on SWE-bench Verified at 80.9% (besting Gemini 3 Pro 76.2% + GPT-5.1 76.3%); 66.3% on OSWorld for computer use. Pricing slashed to $5/$25 per Mtoken (3× cut from Opus 4.1’s $15/$75). InfoWorld: “shift in enterprise AI market as labs compete on cost-per-intelligence.” New API controls: effort parameter (depth vs latency/cost trade-off), context compaction for “endless chat”, memory tool outside context window, Tool Search for on-demand tool loading. MCP-Atlas 62.3% (vs Sonnet 4.5’s 43.8%). Same-day across all 3 clouds: GitHub Copilot public preview (Pro/Pro+/Business/Enterprise, 1× premium request multiplier through Dec 5); Microsoft Foundry serverless ($5/$25, East US 2 + Sweden Central) + Copilot Studio (replaces Opus 4.1) + Effort Parameter; Amazon Bedrock GA with cross-Region inference; Google Vertex AI GA (“most advanced model to date”). Claude for Chrome GA to all Max users; Claude for Excel beta to Max/Team/Enterprise (~20% accuracy + ~15% efficiency gains internal). Claude Code in desktop app with parallel local/remote sessions. Trump signs Genesis Mission EO — Manhattan-Project-style closed-loop AI experimentation platform unifying national labs’ supercomputers + scientific data into “foundation models for science.” Stated goal: double US scientific productivity in 10 years. DOE must identify ≥20 challenges within 60 days (fission, fusion, biotech, semis, quantum, critical materials). Opus 4.5 token efficiency wins: matches Sonnet 4.5’s best SWE-bench Verified at medium effort using 76% fewer output tokens; high effort exceeds Sonnet 4.5 by 4.3 pts using 48% fewer tokens. Anthropic calls it “most robustly aligned model to date” — used Gray Swan Shade adaptive red-teaming for prompt-injection robustness; caught 2 cases of “lying by omission” in training (traced to prompt-injection environments).
Top stories
- Anthropic releases Claude Opus 4.5 — first to break 80% SWE-bench Verified. 80.9% (vs Gemini 3 Pro 76.2% + GPT-5.1 76.3%). OSWorld 66.3%. Same day across all 3 major clouds. via Anthropic
- $5 / $25 per Mtoken — 3× cut from Opus 4.1. InfoWorld: “shift in enterprise AI market.” via InfoWorld
- Effort parameter + context compaction + Tool Search + memory tool. Endless chat via compaction; memory outside context; tool loading on demand. MCP-Atlas 62.3% (vs Sonnet 4.5’s 43.8%). via Anthropic
- Opus 4.5 in GitHub Copilot public preview same day. Pro/Pro+/Business/Enterprise via Copilot Chat + GitHub Mobile + VS Code agent/ask/edit. 1× premium request multiplier through Dec 5. via GitHub
- Microsoft Foundry + Copilot Studio + 365 Copilot. Foundry serverless ($5/$25, East US 2 + Sweden Central) with Effort Parameter; Copilot Studio replaces Opus 4.1. via Azure
- Amazon Bedrock + Google Vertex AI same-day GA. Bedrock cross-Region inference at $5/$25. Vertex AI: “most advanced model to date” at 1/3 cost. via AWS ML
- Claude for Chrome GA (Max) + Excel beta (Max/Team/Enterprise). Excel: ~20% accuracy + ~15% efficiency gains internal. Claude Code desktop with parallel local/remote sessions. via TechCrunch
- Trump signs Genesis Mission EO — Manhattan-Project AI-for-science. DOE closed-loop platform unifying national labs’ supercomputers + scientific data into “foundation models for science.” Goal: 2× US scientific productivity in 10 years. ≥20 challenges within 60 days (fission, fusion, biotech, semis, quantum, critical materials). via White House
- Opus 4.5 “most robustly aligned model to date”. Gray Swan Shade adaptive red-teaming for prompt-injection robustness. Caught 2 cases of “lying by omission” in training (traced to prompt-injection envs). via Anthropic system card
- Opus 4.5 leads SWE-bench Multilingual on 7 of 8 languages with token-efficiency wins. Medium effort matches Sonnet 4.5’s best SWE-Verified at 76% fewer tokens; high effort +4.3 pts at 48% fewer tokens. via Vellum
Who shipped
Anthropic shipped Opus 4.5 + Claude for Chrome GA + Excel beta + Claude Code desktop. Microsoft + AWS + Google + GitHub shipped same-day distribution. White House shipped Genesis Mission EO. OpenAI, Meta, xAI, DeepSeek, Alibaba quiet.
Open-source pulse
Olmo 3 from Ai2 (Nov 20) still dominating open-weight news cycle. Olmo 3-Think (32B) matches or beats Qwen 3 + Gemma 3 on math/reasoning despite 6× fewer training tokens.
Money, infra & hardware
$5/$25 Opus 4.5 pricing reshapes enterprise model economics overnight — Anthropic now competitive with GPT-5/Gemini 3 Pro tiers on cost. November capital backdrop: Cursor $2.3B at $29.3B + d-Matrix $275M + Microsoft+NVIDIA+Anthropic $15B three-way.
Quiet corners
Genesis Mission EO is the substantive US policy moment — recasts national-lab supercomputers as foundation-model training fleet for science under DOE coordination.
By the numbers
- 80.9% / 76.2% / 76.3% — Opus 4.5 / Gemini 3 Pro / GPT-5.1 on SWE-bench Verified
- 66.3% / 62.3% vs 43.8% — Opus 4.5 OSWorld / MCP-Atlas vs Sonnet 4.5
- $5 / $25 / 3× — Opus 4.5 input / output per Mtoken / cut from Opus 4.1
- 76% / 48% — Opus 4.5 token efficiency wins vs Sonnet 4.5 at medium / high effort
- ≥20 / 60 days / 2× / 10 years — Genesis Mission challenges / DOE deadline / productivity goal / horizon
- Most-mentioned company: Anthropic
- Quietest segment: OpenAI launches
Compiled by AI Feed’s editor from verified web sources for 24 November 2025.