Devs – you have 64gb of VRAM – which model do you use for coding?
I've currently settled on an unsloth version of Qwen 3.5 122b-a10b model (UD-IQ4_NL). With 100k bf16 context window, I only had to load a few layers…
I've currently settled on an unsloth version of Qwen 3.5 122b-a10b model (UD-IQ4_NL). With 100k bf16 context window, I only had to load a few layers…
https://www.tomshardware.com/pc-components/dram/meta-fights-soaring-hardware-costs-by-reusing-old-ddr4-server-memory-in-new-ddr5-only-servers-custom-cxl-2-0-chip-marries-legacy-ddr4-2400-with-cutting-edge-ddr5-6400 submitted by /u/pulse77 [link] [comments]
My last post got a lot of interaction asking 6000 pro owners if they regretted, the answer was hard NO. I ended up understanding that dual…
I know people have been talking about creating a “second brain” with local AI trained on personal information, but I’m curious about how that actually played…
I saw someone test Qwen3.6-27B with a 3-critic harness. The harness included code review, test review and Playwright e2e. Each critic had context. The result was…
Overall Performance Gains: Qwen3.5 4B: +36.1% Qwen3.6 27B: +18.9% Gemma4 12B: +65.1% Overall average: ~40% Only for gfx900 related GPUs: Vega GPU, codename vega10, including Radeon…
submitted by /u/arduinoRPi4 [link] [comments]
We kept hitting the same wall building multi-hop RAG: the systems with the best accuracy (GraphRAG, HippoRAG, RAPTOR) all lean on a knowledge graph built offline…
Hey everyone! I open-sourced something I've been working on called Lullabeast. It's an autonomous dev pipeline. You describe your project and planner, executor, and reviewer agents…
Over a year ago, we set out to build a single-turn full-book writing model. Half a year ago, we published our LongPage Dataset for book scale…