activity
20242026
most citedSV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding

1 citations · 3 across the 14 of their papers we have counts for

collaborators
Showing 2024 · cs.CVShow all

7 papers · 2 filters

cs.CV2024

A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation

Shijie Zhou, Ruiyi Zhang, Yufan Zhou +1

Large multimodal models still struggle with text-rich images because of inadequate training data. Self-Instruct provides an annotation-free way for generating instruction data, but…

cs.CV2024★ 1 cited

SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding

Jian Chen, Ruiyi Zhang, Yufan Zhou +6

Multimodal large language models (MLLMs) have recently shown great progress in text-rich image understanding, yet they still struggle with complex, multi-page visually-rich documen…

cs.CV2024

MMR: Evaluating Reading Ability of Large Multimodal Models

Jian Chen, Ruiyi Zhang, Yufan Zhou +3

Large multimodal models (LMMs) have demonstrated impressive capabilities in understanding various types of image, including text-rich images. Most existing text-rich image benchmar…

cs.CV2024

LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models

Ruiyi Zhang, Yufan Zhou, Jian Chen +3

Large multimodal language models have demonstrated impressive capabilities in understanding and manipulating images. However, many of these models struggle with comprehending inten…

cs.CV2024

Diffusion Models For Multi-Modal Generative Modeling

Changyou Chen, Han Ding, Bunyamin Sisman +5

Diffusion-based generative modeling has been achieving state-of-the-art results on various generation tasks. Most diffusion models, however, are limited to a single-generation mode…

cs.CV2024★ 1 cited

TRINS: Towards Multimodal Language Models that Can Read

Ruiyi Zhang, Yanzhe Zhang, Jian Chen +4

Large multimodal language models have shown remarkable proficiency in understanding and editing images. However, a majority of these visually-tuned models struggle to comprehend th…