r/LocalLLaMA
· Communities
Does MTP head get loaded in VRAM by default?
I ran into a doubt when using the following command. It seems that the System RAM usage keeps increasing even though there is >10GB of space left in VRAM while using the MTP mode. Does the MTP head load separately from the main model? Do I need to set the device here as well? /mnt/ml/llama.cpp/llama.cpp-cuda-13.2-20260