Deepseek-V4-Flash-0731 Dwarfstar on Mac
Here is the prefill performance in an M2 Ultra with 192GB of RAM. For decode, at the following depth: Start: 28 t/s 45k: 23.5 t/s 192k:…
Here is the prefill performance in an M2 Ultra with 192GB of RAM. For decode, at the following depth: Start: 28 t/s 45k: 23.5 t/s 192k:…
I have a bunch of PDFs, word docs, and excel files that relate to a project of interest. I am trying to figure out what is…
previously on my 3070, 32gb ddr4 and i711700 I used this command for months and got 26-30 tps: "C:Program Filesllama cppllama-server.exe" ^ -m "C:Program Filesllama cppmodelsQwen3.6-35B-A3B-UD-Q4_K_XL.gguf"…
I let Gemma4-31b run on my laptop for like almost a day using a heavily altered pi to do a deep dive on our beloved Llama…
How is it possible for such an active group like Unsloth to quantize so many models, and yet Hy3, which came out at the start of…
I know qwen 3.5 4b is great but a bit too large and miniPCM5 1b is great for agentic use but not so great for multilingual…
submitted by /u/mailto_devnull [link] [comments]
submitted by /u/rmhubbert [link] [comments]
https://huggingface.co/tsfrm/vacuum-16t A 16.5-trillion-parameter model that contains nothing. This model is just a ████ you to the labs and companies who say that "haha I have the…
We all have been there, tinkering around with models is fun but we rarely do it with research precision and issues are often subtle and hard…