Skip to content
arXiv cs.CV · Papers

Scalable Visual Pretraining for Language Intelligence

arXiv:2607.09657v1 Announce Type: new Abstract: The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information