Unsloth has uploaded several sizes of Deepseek-V4-Flash GGUF’s
submitted by /u/ForsookComparison [link] [comments]
submitted by /u/ForsookComparison [link] [comments]
It's more of a showcase demo at the moment, but Qwen3TTS 0.6B runs usable fast in Q4. On my Galaxy S25 I get around 0.5 realtime…
VisionBridge lets you give text-only LLMs vision. It's tiny OpenAI-compatible proxy that lets reasoning models (DeepSeek, Qwen, GLM…) see images by querying a separate vision model…
We have 8xB200 nodes and users keep asking us how to serve GLM-5.2 on them. Our engineering team went through everything published so far, and the…
submitted by /u/cuolong [link] [comments]
This is specifically for coding, technical planning, and hardware setup. _____ The only times Qwen 3.6 35B A3B has let me down, it has been something…
We just open-sourced Gepard 1.0, a TTS model built for real-time conversation. It’s streaming-first: audio starts the moment text arrives, generated frame by frame instead of…
Hey guys, A month ago I posted my MTP benchmarks here (3.34x on Gemma 4). DFlash support just merged into llama.cpp (PR #22105), so I ran…
Lower is better - Quantization increases from right to left I recently made a post here about how I squeezed more context into a Q8 model…
I keep running into this annoying gap: A model exists on Hugging Face. Sometimes it has GGUFs. It runs fine locally in Ollama / llama.cpp /…