Skip to content
r/LocalLLaMA · Communities

SLMs & QAT

I know many labs trying to shrink deployment costs and increase efficiency. While I do think that that is fine and dandy, I do sometimes question why they bother going so small, and yet training with all 16 bits. Wouldn't it be better for labs behind models like Nanbeige, Liquid and Qwen (when focusing on small models,