Skip to content
r/LocalLLaMA · Communities

Local LLM open-source model options (5060TI 16GB)

Getting the above usage rate from running Qwen2.5-14B with the commands below ./llama-cli -m /home/XXXX/huggfacemodels/Qwen2.5-14B-Instruct-Q4_K_M.gguf -ngl 99 -c 32768 [ Prompt: 667.8 t/s | Generation: 44.0 t/s ] I think i can do better as there are still some headroom available on the gpu/cpu Any better way to get mo