r/LocalLLaMA
· Communities
Need some help figuring which quant sizes of some of the big MoEs would fit properly (with how much context + unquantized KV) on a future 256GB or 512GB dram + 48GB VRAM rig I might build later on, since I want to download and save them now (not later on when I have the rig)
I plan to download a few of them in full 16-bit safetensors (so I'd be able to turn those into whatever quant size/quant-style I want later on), but I don't have enough storage space to just grab all the big models all in 16-bit and store all of them like that, so, I'll probably just get 1 or 2 of the biggest ones in 1