Laguna S 2.1 Thinking mode
If many people have noticed that there's no reasoning phase in Laguna S 2.1. I noticed the Poolside development team updated the chat template twice in…
If many people have noticed that there's no reasoning phase in Laguna S 2.1. I noticed the Poolside development team updated the chat template twice in…
Poolside have updated the full precision, and FP8 versions with a fix for the looping issue many of us have been seeing. Other variants incoming. Discussion…
https://x.com/Badtheorylabs/status/2079306502897074249 A 27B open-weight agent model built for agentic coding, structural tool use . The complete thing fits in one 8.39GB file under 2.5 bits per…
Today the Department of Energy (DOE) and Arcee AI announced the development of Genesis-Science-1 (GS1), an open model for scientific research. This is a joint effort…
Hey r/LocalLLM, Built Project Zero — a from-scratch CPU-only LLM inference engine in pure C99. It beats bitnet.cpp by 1.8× on the same hardware. We also…
Models: (Check Model cards for so much sample demo images) https://huggingface.co/microsoft/Mage-Flow https://huggingface.co/microsoft/Mage-Flow-Turbo https://huggingface.co/microsoft/Mage-Flow-Edit Mage-Flow is a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based…
Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the…
Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty…
submitted by /u/ZenaMeTepe [link] [comments]
GLM-5.2 UD-Q4_K_XL GGUF @ 12.2 tok/s output // 30.9 tok/s input on a real 10.7k-token document using llama.cpp RPC - At 10.7k context: 10.2 tok/s output…