VibeVoice 1.5B Running Locally…On an iPhone! Only ~2.2 GB of Memory and Up to 1.28× Real-Time Speed
I speed up the generation part of the demo in case you get bored 😄 I also tested another long-form generation, and the VRAM usage looks…
I speed up the generation part of the demo in case you get bored 😄 I also tested another long-form generation, and the VRAM usage looks…
So I'm looking at https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF and I realize my 128GB of DRAM just isn't cutting it for this (incredibly powerful) model. If only I had another…
So one of yall mentioned that cuda 13.1 or 13.2 is broken for unsloth so I looked in to it, and they were right. I had…
submitted by /u/cafedude [link] [comments]
https://github.com/yhfgyyf/vllm-deepseek-v4-sm89 I couldn't believe that someone actually got vLLM working with this particular set of GPUs, but here it is. The video is from right after…
Took the REAP adaptation of DeepSeek-V4-Flash (0xSero/DeepSeek-V4-Flash-0731-REAP) along with antirez/deepseek-v4-gguf as inspiration, and decided to see how aggressive we could get with standard quant tricks to…
submitted by /u/realmvp77 [link] [comments]
It is one of the best local models ever released, in both 20B and 120B versions. I always come back to it, especially the 120B version.…
Liquid AI released LFM2.5-2.6B today, and this might be more relevant to local AI than another massive model most people cannot run. The model is only…
More info: https://github.com/lechmazur/debate https://preview.redd.it/5rr60v33bfhh1.png?width=3600&format=png&auto=webp&s=bbe9ff2e3c2e599e887a4ec0da7bcc26f71ad4fb This benchmark measures how well LLMs hold an argument under adversarial, multi-turn opposition across a wide range of topics. It rewards broad…