Skip to content
arXiv cs.CL · Papers

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

arXiv:2608.04010v1 Announce Type: cross Abstract: Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation, incurring substantial memory or latency overhead. More importantly, most existing methods fail to alter the rigid, fixe