Prefill vs. decoding and local LLM ROI: is prefill underrated?
I'm trying to understand why, when people discuss the ROI of running LLMs locally, they almost always focus on output speed (decoding) and rarely on input…
I'm trying to understand why, when people discuss the ROI of running LLMs locally, they almost always focus on output speed (decoding) and rarely on input…
Been a long time lurker of this subreddit, learned a whole lot from here and Gemini. I've finally got my rig somewhere I feel I could…
Looking for some suggestions on what I should be realistically aiming for. Recently I've used qwen 3.6 27b at Q3KM, but I've also been able to…
Just sharing some slop. Used opencode as the harness. I know this model isn't really recommended for coding, but I was just curious how it would…
Got my Ascent GX10 two days ago and spent the last couple of days pushing a REAP-pruned NVFP4 DeepSeek-V4-Flash setup on a single Spark by patching…
Open Computer running in an isolated VM with inference running M4 Pro via LM Studio Gemma 4 13B QAT Hey everyone, Tim from AnythingLLM, where we…
i have an existing system with: Asus ROG MAXIMUS Z790 DARK HERO LGA1700 Intel Core i9-14900K G.Skill Trident Z5 RGB 2x48GB DDR5 6800MHz CL34 (x2, 4…
Weights, all 4 sizes, Apache-2.0 (ViT-S / ViT-B / ViT-L / ViT-g): https://huggingface.co/collections/robbyant/lingbot-vision Code: https://github.com/robbyant/lingbot-vision Project page: https://technology.robbyant.com/lingbot-vision Self-supervised DINO-family backbone, but the masking is boundary-driven:…
We rigorously evaluate the resulting checkpoint across general reasoning, non-reasoning multiple-choice question answering, everyday multi-turn conversations, system prompt adherence, safety, math, code and agentic use cases.…
https://www.gmktec.com/products/gmktec-evo-x3-ai-mini-pc-amd-ryzen-ai-max-395 Features: USB 4. Adds a dedicated OCuLink port. OCuLink provides a direct, high-speed cable connection to a desktop graphics card, minimizing data loss and improving…