3 papers
cs.CV2024
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
Shijie Zhou, Ruiyi Zhang, Yufan Zhou +1
Large multimodal models still struggle with text-rich images because of inadequate training data. Self-Instruct provides an annotation-free way for generating instruction data, but…
cs.CV2024
MMR: Evaluating Reading Ability of Large Multimodal Models
Jian Chen, Ruiyi Zhang, Yufan Zhou +3
Large multimodal models (LMMs) have demonstrated impressive capabilities in understanding various types of image, including text-rich images. Most existing text-rich image benchmar…
cs.CV2024
LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models
Ruiyi Zhang, Yufan Zhou, Jian Chen +3
Large multimodal language models have demonstrated impressive capabilities in understanding and manipulating images. However, many of these models struggle with comprehending inten…