Skip to content
Source · Daily Brief

AI Daily Brief — 2 August 2026

A quiet stretch on the AI calendar: much of the day’s volume came from community benchmark chatter around DeepSeek V4 Flash 0731 rather than fresh headline launches. Even so, a handful of genuine ships landed — Sakana AI’s Japanese API, Thinking Machines Lab’s open-weight Inkling-Small and OpenAI’s enterprise agent push.

Top stories

  • Sakana AI opens its Japanese-specialised Namazu API. A new LLM endpoint tuned for Japanese-language work. via Sakana AI
  • Thinking Machines Lab releases Inkling-Small. A 276B-total, 12B-active open-weights multimodal MoE model. via MarkTechPost
  • OpenAI Presence targets production-ready business agents. A push to make AI agents dependable enough for enterprise deployment. via THE DECODER
  • Claude Opus 5 pushes prompt-to-game generation further. Coverage notes a jump from rough color blocks to 3D prototypes with physics and music. via THE DECODER
  • Interconnects surveys the open-weight frontier. The latest open-artifacts roundup weighs Laguna S2.1, Inkling and Kimi K3 on the Pareto curve. via Interconnects (Nathan Lambert)
  • Grok adds any-video analysis. The assistant can now describe footage frame by frame and flag AI-generated clips. via X · @elonmusk
  • EU rules on AI models become enforceable. New obligations for model providers take effect across the bloc. via Hacker News (front page)
  • Sam Altman and the deceleration debate. A look at the widening argument over how fast the field should move. via TechCrunch – AI

Who shipped

Sakana AI opened access to Namazu, a Japanese-specialised LLM API. Thinking Machines Lab put out Inkling-Small, a 276B-total, 12B-active open-weight multimodal MoE. OpenAI introduced Presence for enterprise agents, and a claimed internal Astra model reportedly cracked several open math and CS problems. Anthropic’s Claude Opus 5 drew coverage for turning prompts into 3D game prototypes. xAI extended Grok with any-video analysis, while DeepSeek and MiniMax dominated the open-weight conversation with V4 Flash and H3.

Open-source pulse

The open-weight side did most of the talking. DeepSeek V4 Flash 0731 was everywhere in the local-model community — KV-cache and quantisation notes, prefill-speed gains, chess and MMLU-Pro benchmark posts, and fresh llama.cpp MTP support. MiniMax H3 shipped open weights the same day, Moonshot’s Kimi K3 kept surfacing in efficiency tests, and Qwen variants like a native-MLX WinterMix build made the rounds. Thinking Machines’ Inkling-Small and poolside’s Laguna S2.1 rounded out a crowded frontier.

Quiet corners

The Western frontier labs were largely off-stage. Google DeepMind, Meta, Mistral and Cohere shipped nothing of note, leaving the day’s model news to Chinese labs and smaller independents.

Money, infra & hardware

Hardware chatter leaned Chinese and local-first. A reported DFSX accelerator claims twice the memory bandwidth of NVIDIA’s GB200 via r/LocalLLaMA, while an analysis pegged Kimi K3 on AMD’s MI355X at better performance-per-dollar than the B300 via Hacker News (front page). On the software side, llama.cpp shipped roughly a dozen tagged builds and Simon Willison released condense-json 1.0, with no major funding rounds surfacing.

By the numbers

  • 130 stories across 23 sources
  • Most-mentioned model: DeepSeek
  • Most-mentioned lab: OpenAI
  • Notable absences: Google, Meta

Compiled by AI Feed’s editor from all 130 headlines published on 2 August 2026.