What happened to Petals (Decentralized Inference) by BigScience?
https://www.reddit.com/r/LocalLLaMA/s/3DAhg3HPIa submitted by /u/-OpenSourcer [link] [comments]
https://www.reddit.com/r/LocalLLaMA/s/3DAhg3HPIa submitted by /u/-OpenSourcer [link] [comments]
The PR : https://github.com/ggml-org/llama.cpp/pull/24162 All to git pull, cmake , and download GGUFs ! A vos marques, prêt, partez ! submitted by /u/Squik67 [link] [comments]
I improved my qwen3-tts.cpp implementation to be about 5x realtime on my RTX 5080. It is GGML based, so it should compile and run anywhere -…
submitted by /u/johnnyApplePRNG [link] [comments]
For those who want to be as paranoid and maximally doomsday prepped as possible, I am curious what the most thorough "doomsday kit" is of things…
submitted by /u/johnnyApplePRNG [link] [comments]
submitted by /u/johnnyApplePRNG [link] [comments]
Identical custom quantization recipe on HuiHui's and Vanilla 3.6-35B-a3B. Somehow removing refusal get's you closer to truth and wisdom (?) in math and coding. Benchmarks on…
All hail Z. Ai submitted by /u/Independent-Wind4462 [link] [comments]
I was expecting what when doubling my VRAM from 24gb to 2x24gb I'd use higher quants with more context, and thus get smarter LLMs, but that's…