Sorry, but did Dario just say that closed-weights, in-secret models are worse than open-weights ones?
submitted by /u/BritishDudeGuy [link] [comments]
submitted by /u/BritishDudeGuy [link] [comments]
I was talking to my friend the other day, who is a really avid supporter of local models. He thinks we might be able to get…
submitted by /u/rm-rf-rm [link] [comments]
From the description: "Reasoning-Medical-27B is designed for universal advanced medical reasoning in professional medicine, medical genetics, college biology/medicine, and clinical knowledge. The model was fine-tuned on…
Following some recent commits in llama.cpp, preserve_thinking behavior for chat templates included in older DSV4 ggufs got broken. This makes the model pretty dumb in a…
Most quantization works like this: pick a bit depth, apply it everywhere, maybe let imatrix take a rough guess at what matters, ship it. Most don't…
VibeVoice-ASR-BitNet is a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs — no GPU required. Through heterogeneous quantization, the model is compressed from…
submitted by /u/HugoCortell [link] [comments]
It has only been two days since I move 100% from Qwen3.5-27B F16 to ThinkingCap-Qwen3.6-27B F16. Where I was getting tps in 30-40 range (depending on…
submitted by /u/fulgencio_batista [link] [comments]