Skip to content
arXiv cs.CL · Papers

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

arXiv:2608.00013v1 Announce Type: new Abstract: Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: compute-based scaling laws fail to generalize across model families, and no framework exists f