r/LocalLLaMA
· Communities
Reducing the model parameter size?
If the new releases of Chinese open weight models arrive at 2T+ sizes, is it possible for a research institution (with GPU clusters) to somewhat easily reduce them to smaller models that fit on consumer GPUs, or is it something reasonably feasible only for the original vendor? submitted by /u/tt23 [link] [comments]