3 papers
cs.CV2026
CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
Xuehang Guo, Pingyue Zhang, Ruiyi Zhang +6
Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate…
cs.CV2026
EMCompress: Video-LLMs with Endomorphic Multimodal Compression
Zheyu Fan, Jiateng Liu, Yuji Zhang +4
Video-LLMs face a fundamental tension in long-video reasoning: static, sparse frame sampling either dilutes evidence across task-irrelevant segments at significant cost or misses f…
cs.AI2025
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Rui Yang, Hanyang Chen, Junyu Zhang +10
Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have…