DeepSeek-V4-Flash-0731-UD-Q3_K_XL 3×3090 test results
For anyone interested, here are the llama-bench results on 3 bit K_XL quantization. I think this could be pushed further but no luck so far. Command…
For anyone interested, here are the llama-bench results on 3 bit K_XL quantization. I think this could be pushed further but no luck so far. Command…
https://preview.redd.it/u3xo0qklksgh1.png?width=680&format=png&auto=webp&s=9d72c80bbad2559e210897e137d30c88faa7bbc7 Hehe. submitted by /u/laterbreh [link] [comments]
Thanks to the community help I finally launched this llm. LM Studio refused to load weight onto second GPU but Unsloth Studio did so everything was…
Basically you now have to mark all AI generated images, audio, video and text as AI generated. :P submitted by /u/xoxaxo [link] [comments]
I think people are sleeping on Gemma and local models so I built a free, very fast harness for Gemma 4 that I call Tomte. https://tomteapp.com…
The weights have now been added to the repo an hour ago. This model is built upon LongCat-Flash-Lite, the differences are that LongCat-Flash-Lite-Sparse: Replaces dense MLA…
Personally, i think good coding model shouldn't be focused on one-shot "everything in one html-file" tests, but should be really good on debugging, fixing and modifying…
https://preview.redd.it/eqqebec92sgh1.png?width=873&format=png&auto=webp&s=4e9a120e439dc0a64ee4aef87d84b45c51f6bf02 - 3x32 DDR4 2400 Mhz ECC. CPU itself supports quad channel, and we already know why am i haven't filling those slot yet. - E5…
Poolside have updated the FP8 and NVFP4 checkpoints for Laguna S 2.1, increasing the default context size to 1 million, and updating the configs. Here's hoping…
I've been trying to find a good model to run locally, and in the benchmarks I can Gemma4 does well. However, whenever I give it anything…