109 citations · 201 across the 42 of their papers we have counts for
34 papers · 1 filter
World Tokens: Enhancing Embodied Policies with Training-Time World Modeling
Qu Tang, Benhui Zhuang, Bo Yuan +3
Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes…
TimeThink: Reasoning with Time for Video LLMs
Handong Li, Longteng Guo, Zikang Liu +8
Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Video-LLMs) have shown promisi…
LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video
Shiqiang Lang, Jing Liu, Haoyang He +6
Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon tasks such as autonomous driving…
Semantic-Enriched Latent Visual Reasoning
Tianrun Xu, Yue Sun, Qixun Wang +8
Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. However, existing approaches larg…
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
Longteng Guo, Yifan Wang, Pengkang Huo +4
Recent multimodal large language models (MLLMs) achieve strong performance on visual reasoning benchmarks, yet it remains unclear to what extent such performance reflects reasoning…
SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation
Longteng Guo, Xuanxu Lin, Dongze Hao +5
Scientific reasoning is a key aspect of human intelligence, requiring the integration of multimodal inputs, domain expertise, and multi-step inference across various subjects. Exis…