3 papers
cs.CV2025
A Token-level Text Image Foundation Model for Document Understanding
Tongkun Guan, Zining Wang, Pei Fu +9
In recent years, general visual foundation models (VFMs) have witnessed increasing adoption, particularly as image encoders for popular multi-modal large language models (MLLMs). H…
cs.CV2025
High-Resolution Image Synthesis via Next-Token Prediction
Dengsheng Chen, Jie Hu, Tiezhu Yue +2
Recently, autoregressive models have demonstrated remarkable performance in class-conditional image generation. However, the application of next-token prediction to high-resolution…
cs.CV2024
CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder
Lichen Ma, Tiezhu Yue, Pei Fu +4
Recently, significant advancements have been made in diffusion-based visual text generation models. Although the effectiveness of these methods in visual text rendering is rapidly…