llamacpp patch – DeepSeek V4 Flash running with full 1M token context locally on RTX 5090
Wanted to try running DeepSeek V4 Flash locally but found it asking for absurd amounts of VRAM at higher context lengths (~256GB at 1M). Turned out…
Every primary-source story across every tracked model. Filter by clicking a chip.
Wanted to try running DeepSeek V4 Flash locally but found it asking for absurd amounts of VRAM at higher context lengths (~256GB at 1M). Turned out…
RT Mike Bradley“Within the next 18 months, you will be able to host GLM 5.2 equivalent intelligence on an RTX 5090 GPU.”-Ahmad Osman, AI World’s Fair
We will deprecate Gemini 2.5 Pro and Gemini 3 Flash across all GitHub Copilot experiences (including Copilot Chat, inline edits, ask and agent modes, and code…
RT davincipre-chatgpt openai was a lab. pre-gemini deepmind as well (still somewhat is, maybe?). anthropic almost never was (it's an extremely product-oriented company with very little…
Hey folks. I've been frustrated by how difficult it is to get an idea of how good each new model (or fine-tune) is, and I've not…
So apparently Gemini Omni Flash is to Seedance 2.0 what Seedance 2.0 is to Veo 3. But Seedance 2.0 curb stomped Veo 3… is this real?…
RT Luma | Dream LabRe It started with a strange idea: a planet that was alive.Under constant threat, she forms a covenant with a mana-wielding civilization.…
V4 is really bad at in-context learning, alasMerely GLM 5.0 levelwe shall see how much they can improveDeyao Zhu: [6/n] We evaluated model releases from September…
submitted by /u/9gxa05s8fa8sh [link] [comments]
RT X FreezeFinally, Grok’s Speech-to-Text is now live in Grok BuildYou can now just dictate prompts directly to your coding agents using /voice or Ctrl +…
RT Design ArenaBREAKING: Gemini Omni Flash by @GoogleDeepMind is 1st overall on Video Arena with an Elo of 1404.Gemini Omni Flash establishes a 101 point Elo…
I've been a lurker for a while and have been building my own home lab with P40's and MI50's. I've learned so much from the community…
Article URL: https://github.com/anthropics/claude-code/issues/73125 Comments URL: https://news.ycombinator.com/item?id=48765630 Points: 23 # Comments: 19
The news comes about a week after OpenAI announced its own custom AI chip in a partnership with Broadcom.
I HATE CLAUDE OPUS 4.8 AND I HATE DARIO AMODEIAnd his newest models suggest that the feeling's mutualgum: kimi-k2.7-code scored 75.92% standard and 72.58% strict. glm-5.2…
RT Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)Re @meowbooksj the design space of sparsity is barely exploredLike, why don't we use MLP routers in MoEs? Because le…
We need more of this, 100+ T/s on dense models is the difference between defaulting to Claude/Codex for everything vs having a local private model doing…
I'm struggling to find the right setup with llama.cpp. Ubuntu 26.04, AMD 5900X, 128GB DDR4-3600, R9700s are running PCIe x8/Gen4 Model config: [Qwen3.6-27B] mmproj = /models/Qwen3.6-27B-mmproj-BF16.gguf…
Grok Build is now installed in Railway sandboxesRailway: Grok Build from @xai is now available in Railway sandboxesRun `ssh sandboxes@railway.new` in your terminal and try it…
RT sam lessin 🏴☠️The Narrative Strategy of Sam Altman on CNBC today…Some notes from discussing this ‘we will give America 5% of OpenAI’ idea…1 - Strategically…
Anthropic is reportedly in talks with Samsung Electronics about manufacturing a custom AI chip. The project is still early, but Anthropic has already hired chip engineers.…
this was a long time coming, but it's finally here! you can now basically supercharge whichever UI you're already using with the power of openlumara. click…
Halfway to its Qwen 3.6 peer (I choose to ignore Terminal Bench). Presumably they, too, have learned the Dao of continuous RL gains. If this is…
Fable in Claude Code is capable of really amazing things, including for non-coders, but the interface is not really designed for managing 5+ hour long autonomous…
We track 28 AI models across text, image, video, and audio domains. Each model has its own filtered news feed — click a chip above to see only that model's primary-source coverage. Tagging is automatic at ingest time using a strict title-only keyword match (so a benchmark post that mentions five models in its summary only shows up under whichever model is named in the headline — no thin-content drift).
Text / LLMs: GPT, Claude, Gemini, Gemma, Llama, Mistral, Grok, Qwen, DeepSeek, Kimi, GLM, MiniMax, Yi, Hunyuan, Command, Phi.
Image generation: FLUX, Stable Diffusion, Midjourney, Imagen.
Video generation: Sora, Veo, Runway, Luma, Kling, Pika.
Audio / voice: ElevenLabs, Suno.