r/LocalLLaMA
· Communities
Qwen 30b MoE – 30tps – 6GB vram – Done!
So, I have been dreaming of getting 17 tokens per second using my RTX 3050 6GB version on a decent context window for Hermes needed above 60k. The hope is that has was a 22GB of DDR 4, hoping they can take some of those experts and give me room for context. What did I get 10 or less tokens per second. 😄 Not today!! Tod