3 papers
cs.CV2026
Culture In a Frame: CB as a Comic-Based Benchmark for Multimodal Culturally Awareness
Yuchen Song, Andong Chen, Wenxin Zhu +4
Cultural awareness capabilities have emerged as a critical capability for Multimodal Large Language Models (MLLMs). However, current benchmarks lack progressed difficulty in their…
cs.AI2026
Thinking with Comics: Enhancing Multimodal Reasoning through Structured Visual Storytelling
Andong Chen, Wenxin Zhu, Qiuyu Ding +3
Chain-of-Thought reasoning has driven large language models to extend from thinking with text to thinking with images and videos. However, different modalities still have clear lim…
cs.CL2025
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
Wenxin Zhu, Andong Chen, Yuchen Song +4
With the remarkable success of Multimodal Large Language Models (MLLMs) in perception tasks, enhancing their complex reasoning capabilities has emerged as a critical research focus…