A q2_k_s walks into a bar and
sat on a mat. submitted by /u/Risen_from_ash [link] [comments]
sat on a mat. submitted by /u/Risen_from_ash [link] [comments]
I'm constantly bombarded by non-local LLMs in this sub but god forbid I post a local model meme. submitted by /u/fragment_me [link] [comments]
What speeds are everyone getting with deepseek v4 flash 0731? I’m getting~200 tps prompt processing / ~11 tps token gen, on 4x5060ti16gb with ddr4 3200 ram…
I mean seriously y’all, what an amazing past few days. So many awesome new models to test out in the mid range model sizes. submitted by…
Antirez stealthily uploaded the new weights in the old folder... and there we were tapping our fingers. https://huggingface.co/antirez/deepseek-v4-gguf/tree/main submitted by /u/challis88ocarina [link] [comments]
audio.cpp 0.5 is out :) The most fun new model in 0.5 is DramaBox. It is closer to prompt-directed voice acting. DramaBox is built on the…
I'm an avid user of Deepseek v4 Flash via antirez's DS4 DwarfStar inference engine, and so when the new checkpoint dropped, the first thing I did…
Laguna 2.1 at NVFP4 Deepseek v4 at Q2 Inkling-Small at IQ3 Which models you guys running now ? How it compares to 122b? Upcoming in few…
I ran a pre-registered ablation on a classification task (Kubernetes issue → SIG triage) using a 4B model on a 6GB laptop GPU. Same frozen weights,…
I can run the mediums, but sometimes I want a faster option that's smarter than Qwen 27B/35B. On my hardware I get like 500 to 800…