Hypothetically speaking…
Would it not be possible to create crowd sourced, truly open sourced distilled LLMs with a simple wrapper around command line based AI services that exist…
Would it not be possible to create crowd sourced, truly open sourced distilled LLMs with a simple wrapper around command line based AI services that exist…
DeepSpec DeepSpec is a full-stack codebase for training and evaluating draft models for speculative decoding. It contains data preparation utilities, draft model implementations, training code, and…
submitted by /u/sammcj [link] [comments]
TLDR: Biggest quant that fits 100% in VRAM beats a higher quant that spills. And if your MoE ships an MTP draft head, the draft can't…
With latest trends, I can't help but feel squeezed in this current situation, and fear for the worst soon. Open weights for 96GB to 128GB hardware…
Hi, I have 768gb ddr5 6400 ecc ram, given current ram prices should I sell half and buy rtx 6000 pros? submitted by /u/No-Paper-557 [link] [comments]
Sunday experiment. Same prompt to both. Build a voxel world in plain C. No engine, no game library, no framework, just the compiler. The model does…
I'm building a local LLM workstation and would appreciate some advice from people already running 2×3090s. Current hardware: ASUS Crosshair VIII Hero (X570) One Gainward Phoenix…
hey, I’m currently getting enough VRAM to run something in the GLM-5.2 range, but I’m wondering: do we actually have a solid ranking that compares closed-source…
Sharing popular(also recent) models for reference: 151-250B : DeepSeek-V4-Flash Step-3.X-Flash Command-a-plus-05-2026 Laguna-M.1 MiniMax-M2.X Qwen3-235B-A22B 100-150B : GLM-4.5-Air Qwen3.5-122B-A10B NVIDIA-Nemotron-3-Super-120B-A12B Mistral-Small-4-119B-2603 Devstral-2-123B-Instruct-2512 Mistral-Medium-3.5-128B Llama-4