How i got Bonsai-Ternary-27B to run at 120k context <10gb vram.
Hey guys, so its quite a read, its long. If you are bored just read the TLDR and see if its worth it for you: TLDR;…
Hey guys, so its quite a read, its long. If you are bored just read the TLDR and see if its worth it for you: TLDR;…
Article URL: https://fortune.com/2026/07/14/data-centers-23-billion-electricity-bills/ Comments URL: https://news.ycombinator.com/item?id=48914683 Points: 36 # Comments: 9
Article URL: https://www.starfleetmath.com/ Comments URL: https://news.ycombinator.com/item?id=48914646 Points: 13 # Comments: 2
Article URL: https://vpd.ca/ Comments URL: https://news.ycombinator.com/item?id=48914644 Points: 6 # Comments: 4
I still sometimes see people saying "if you know how to write the code, it's faster to write it yourself"I'd argue the exact opposite: if you…
My first time doing anything like this, and I built it because I wanted it. If anyone wants the non Heretic I'll mosey that out as…
RT dexRe great ep @DavidOndrej1 @swyx https://www.youtube.com/watch?v=EWk9PBbKqzc
RT tetsuoGrok 4.5 Usage reset to 0% time to build.
audio.cpp again. Hopefully you are not sick of it yet :) Release 0.3 adds five new models: Supertonic 3, MOSS-TTS-Local, MOSS-TTS-Nano, IndexTTS2, and Irodori-TTS. The highlight…
Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language model (LLM) agents can explore effectively when interacting with one…
June 2026 is about visibility and trust with a clearer view of your GitHub Copilot usage, a new trust layer for MCP servers, and the first…
Together AI offers day zero access to Inkling, Thinking Machines Lab's multimodal mixture-of-experts model for text, image, and audio reasoning.
Large Language Models (LLMs) are increasingly deployed to autonomously solve real-world tasks. A key ingredient for this is the LLM Function-Calling paradigm, a widely used approach…
Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge but still suffers from long contexts and disjoint retrieval–generation optimization. In this work, we propose…
Working at the frontier: Why Base44 trusts Claude Fable 5 with their most challenging engineering work
See how Together AI is improving production GPU clusters with passive health checks, node repair, stronger Slurm reliability, OIDC, and startup scripts.
Visual generative models (e.g., diffusion models) typically operate in compressed latent spaces to balance training efficiency and sample quality. In parallel, there has been growing interest…
Can a kite find freedom after losing its string? 🪁Discover “Kite” by Shihua Lin, Mingwei Chen & Xuanwei Liu.Congratulations on winning the Apex Award at Kling…
a continuation: Codex adding 1M users a day now.
RT Haiyu WuLocally straightening the latent trajectory can largely reduce the optimization difficulty of world model planning!Sounds confusing? Not sure how to achieve it? Don't worry.After…
This update brings major advances in customization and model provider flexibility to all tiers of GitHub Copilot for JetBrains IDEs. With richer plugin and provider experiences,…
PrismML ships Bonsai 27B on a phone, xAI's Grok Build found uploading codebases, Meta tops the Physics Olympiad and faces two AI lawsuits, and Hassabis calls…
RT Latent.Space🆕 5 Trends That Defined AI Engineering at World’s Fair 2026https://latent.space/p/aiewf26trends@ricmac's big recap of @aidotengineer:1. The focus shifts from agents to systems2. Loop engineering is…
At this year's AIE World’s Fair, AI engineering entered a new phase: building systems around agents, rather than just building with agents.