9 papers
CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
Xuehang Guo, Pingyue Zhang, Ruiyi Zhang +6
Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate…
MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing
Zihao Lin, Wanrong Zhu, Jiuxiang Gu +8
Real-world design documents (e.g., posters) are inherently multi-layered, combining decoration, text, and images. Editing them from natural-language instructions requires fine-grai…
OIDA-QA: A Multimodal Benchmark for Analyzing the Opioid Industry Documents Archive
Xuan Shen, Brian Wingenroth, Zichao Wang +12
The opioid crisis represents a significant moment in public health that reveals systemic shortcomings across regulatory systems, healthcare practices, corporate governance, and pub…
SOHES: Self-supervised Open-world Hierarchical Entity Segmentation
Shengcao Cao, Jiuxiang Gu, Jason Kuen +7
Open-world entity segmentation, as an emerging computer vision task, aims at segmenting entities in images without being restricted by pre-defined classes, offering impressive gene…
Towards Visual Text Grounding of Multimodal Large Language Model
Ming Li, Ruiyi Zhang, Jian Chen +7
Despite the existing evolution of Multimodal Large Language Models (MLLMs), a non-neglectable limitation remains in their struggle with visual text grounding, especially in text-ri…
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
Lin Zhang, Zefan Cai, Yufan Zhou +10
Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manua…