collaborators

5 papers

cs.AI2026

MentalThink: Shaping Thoughts in Mental SVG World

Kangheng Lin, Jisheng Yin, Dingming Li +11

We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink…

cs.CV2026

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

Yana Wei, Hongbo Peng, Yanlin Lai +14

We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness. Shifting evaluation from h…

cs.DC2024

DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models

Zili Zhang, Yinmin Zhong, Yimin Jiang +6

Multimodal large language models (LLMs) empower LLMs to ingest inputs and generate outputs in multiple forms, such as text, image, and audio. However, the integration of multiple m…

cs.CV2024

Focus Anywhere for Fine-grained Multi-page Document Understanding

Chenglong Liu, Haoran Wei, Jinyue Chen +7

Modern LVLMs still struggle to achieve fine-grained document understanding, such as OCR/translation/caption for regions of interest to the user, tasks that require the context of t…

cs.CV2024

OneChart: Purify the Chart Structural Extraction via One Auxiliary Token

Jinyue Chen, Lingyu Kong, Haoran Wei +6

Chart parsing poses a significant challenge due to the diversity of styles, values, texts, and so forth. Even advanced large vision-language models (LVLMs) with billions of paramet…