How much are RTX PRO 6000s going for in your country/state?
Hello guys, hoping you're doing fine! On the last 2-3 months, price of the RTX 6000 PRO seem to have gone insane. I will start on…
Hello guys, hoping you're doing fine! On the last 2-3 months, price of the RTX 6000 PRO seem to have gone insane. I will start on…
Gemma 4 was updated (mostly chat templates) and I took it for a test. On a local llama.cpp server running on M5 Pro with 48GB, 26B…
And gathered a lot of data. you can see them for yourself And For the most curious, there are additional details here In this graph, I…
Shoutout to this awesome guy - https://www.reddit.com/r/LLM/s/IDUyU3v9ap Thanks to his project, BigMoeOnEdge https://github.com/Helldez/BigMoeOnEdge, I managed to successfully run a 35B MoE model on just 12GB of…
If you're running Laguna S 2.1 on llama.cpp and hitting thinking loops because it won't close its tags, you might want to look at your quant…
Original post: https://www.reddit.com/r/LocalLLaMA/comments/1tpdt5m/behold_probably_the_most_ghetto_local_ai_server/ I promised a writeup, but didn't have time yet, sorry. I barely had time to do this controller. submitted by /u/MackThax [link] [comments]
TLDR: I (with the help of AI) re-implemented every Blackwell-only kernel (DeepGEMM, FlashInfer sparse-MLA, block-scaled FP8) in Triton, because they simply don't exist for sm89. The…
/s submitted by /u/Porespellar [link] [comments]
Now live on OpenRouter, and free to use through August 3, 2026. Hoping they will going openweight soon~ submitted by /u/niacolhealth [link] [comments]
https://openrouter.ai/inclusionai/ling-3.0-flash submitted by /u/derspenti [link] [comments]