4 papers · 1 filter
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
Jovana Kondic, Pengyuan Li, Dhiraj Joshi +24
Understanding charts requires models to jointly reason over geometric visual patterns, structured numerical data, and natural language -- a capability where current vision-language…
LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning
Haiying Xu, Zihan Wang, Song Dai +3
Despite recent advances in multimodal reasoning, representing auxiliary geometric constructions remains a fundamental challenge for multimodal large language models (MLLMs). Such c…
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
Harold Haodong Chen, Disen Lan, Wen-Jie Shu +10
The rapid evolution of video generative models has shifted their focus from producing visually plausible outputs to tackling tasks requiring physical plausibility and logical consi…
Video Finetuning Improves Reasoning Between Frames
Ruiqi Yang, Tian Yun, Zihan Wang +1
Multimodal large language models (LLMs) have made rapid progress in visual understanding, yet their extension from images to videos often reduces to a naive concatenation of frame…