2 papers
cs.CV2025
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
Jingqun Tang, Chunhui Lin, Zhen Zhao +15
Text-centric visual question answering (VQA) has made great strides with the development of Multimodal Large Language Models (MLLMs), yet open-source models still fall short of lea…
cs.CV2024
Harmonizing Visual Text Comprehension and Generation
Zhen Zhao, Jingqun Tang, Binghong Wu +7
In this work, we present TextHarmony, a unified and versatile multimodal generative model proficient in comprehending and generating visual text. Simultaneously generating images a…