i tried ternary decomposition instead of quantization. it works as good at q4km but takes slightly more vram. while being completly ternary. and completly PTQ (no QAT)
https://arxiv.org/pdf/2607.13511 submitted by /u/LMTLS5 [link] [comments]
https://arxiv.org/pdf/2607.13511 submitted by /u/LMTLS5 [link] [comments]
I figure I can run 30b or even 72b models on it, but this is my first time running my own local LLM. I want to…
Title do you think its likely this small model class will continue to improve at this speed. i'm running q4 unsloth quant of gemma 4 4b…
This is a full teardown video of the NVIDIA H200 NVL and installation of an EK-Pro H200 NVL Water Block, covering the disassembly, prep and mount…
I love Obsidian, but I always wanted to actually talk to my vault, ask questions across my own notes and docs. I just didn't want to…
submitted by /u/pscoutou [link] [comments]
benchmarks as of july 16 2026 -- Apache 2.0 Thinking Machines Lab: Inkling, July 15, 2026 -- MIT DeepSeek V4 Pro, April 24, 2026 Xiaomi MiMo-V2.5-Pro,…
So, I finished reading this paper: Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour The What? This paper, takes on the problem of scaling experiments…
Specifically Qwen3.6-35B-A3B-uncensored-heretic-Q8_0.gguf, temp 0.0, "Create an SVG of a Darth Vader." (In general text use, anything 4 experts or less seems to seriously break down. (8…
For those who look at tool use across multiple runs. I am personally quite fond of sankey charts. submitted by /u/SnooPeripherals5313 [link] [comments]