1 paper
Junpeng Liu, Tianyue Ou, Yifan Song +6
Text-rich visual understanding-the ability to process environments where dense textual content is integrated with visuals-is crucial for multimodal large language models (MLLMs) to…