Is turboquant any good?
I know im late to the party. I was thinking since some time has passed, has turboquant matured enough to be used? Do any of you…
I know im late to the party. I was thinking since some time has passed, has turboquant matured enough to be used? Do any of you…
Hi r/LocalLLaMA — I’m sharing an experimental GPU-only inference backend and looking for independent reproductions, not just stars. Model: tiiuae/Falcon3-10B-Instruct-1.58bit GPU: NVIDIA RTX 5070 Batch: 1…
TL;DR llama.cpp fork with more KV cache quantization features, with all claims supported by benchmarks: KVarN, KV cache precision tail, additional types of standard KV cache…
Last week someone here said ThinkingCap and Fable Fusion "really do beat the OG" for agentic work, so I ran it: 6 self-grading tasks, 5 reps,…
LTT Labs recently received the Linux version of the AMD Ryzen AI Halo for testing, but it turns out that AMD had intended to send the…
https://x.com/i/status/2081398564345802934 submitted by /u/jacek2023 [link] [comments]
Title, for local inference and training, please. submitted by /u/Desperate_Tea304 [link] [comments]
Hey r/LocalLLaMA — maintainer here, obviously biased. OpenSmith is an open-source Python tracing tool for LLM pipelines. The idea: drop u/trace on any function, run opensmith…
The server rack was humming at maximum capacity, fan speeds pegged at 100%, and thermal throttling was the absolute last thing on the Router’s mind. "Give…
submitted by /u/pscoutou [link] [comments]