MiniMax (official) on X: "Open weights. Open research. Open innovation.🫶 Marching for an open future.🤍
submitted by /u/RhubarbSimilar1683 [link] [comments]
submitted by /u/RhubarbSimilar1683 [link] [comments]
Tell me, to get on 20k context and ingestion 44tks, generation 8tks is good numbers for 4x 8880 v4 cpus, 1tb 32channels ddr3 ram and 2x…
submitted by /u/Time_Reaper [link] [comments]
I know im late to the party. I was thinking since some time has passed, has turboquant matured enough to be used? Do any of you…
Hi r/LocalLLaMA — I’m sharing an experimental GPU-only inference backend and looking for independent reproductions, not just stars. Model: tiiuae/Falcon3-10B-Instruct-1.58bit GPU: NVIDIA RTX 5070 Batch: 1…
TL;DR llama.cpp fork with more KV cache quantization features, with all claims supported by benchmarks: KVarN, KV cache precision tail, additional types of standard KV cache…
Last week someone here said ThinkingCap and Fable Fusion "really do beat the OG" for agentic work, so I ran it: 6 self-grading tasks, 5 reps,…
LTT Labs recently received the Linux version of the AMD Ryzen AI Halo for testing, but it turns out that AMD had intended to send the…
https://x.com/i/status/2081398564345802934 submitted by /u/jacek2023 [link] [comments]
Title, for local inference and training, please. submitted by /u/Desperate_Tea304 [link] [comments]