2 papers
cs.CL2024
Autoregressive Pre-Training on Pixels and Texts
Yekun Chai, Qingyi Liu, Jingwu Xiao +3
The integration of visual and textual information represents a promising direction in the advancement of language models. In this paper, we explore the dual modality of language--b…
cs.CL2024
On Training Data Influence of GPT Models
Yekun Chai, Qingyi Liu, Shuohuan Wang +3
Amidst the rapid advancements in generative language models, the investigation of how training data shapes the performance of GPT models is still emerging. This paper presents GPTf…