Scalable Visual Pretraining for Language Intelligence
Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen
Read on arXiv →Key claim
Visual pretraining outperforms text-only methods for language models.
In plain English
Many large language models rely solely on text for training, which overlooks valuable visual information found in documents and web pages. Current methods convert these rich visual sources into plain text, losing important context. This paper proposes a new approach that utilizes visual pretraining directly from these documents, showing that it consistently outperforms text-only pretraining. Builders should care because this could lead to more effective models that better understand and utilize visual data.
The paper introduces a new paradigm of visual pretraining that challenges the text-only approach.
The study is systematic and shows consistent improvements across multiple benchmarks.
Deep reliability assessment
The methodology supports the claim that visual pretraining can outperform text-only pretraining on the same corpus, but it may overclaim the generalizability of these results across all types of language intelligence tasks without further evidence.
Reproducibility
No open source code or dataset is mentioned in the paper.
Key figure
Figure 1 illustrates the matched text and visual pretraining pathways from the same scientific-document corpus, highlighting the differences in how page-level structures are processed.
