arXiv cs.CV
· Papers
Scalable Visual Pretraining for Language Intelligence
arXiv:2607.09657v1 Announce Type: new Abstract: The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information