From the 1 of 9 linked papers with an AI index.
9 papers
Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
Tianyi Gao, Han Fang, Tianyi Ding +9
Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing…
MoAKE: Toward Unified All-in-One Action Quality Assessment via Mixture of Action Knowledge Experts
Huangbiao Xu, Huanqi Wu, Xiao Ke +3
Action Quality Assessment (AQA) aims to objectively evaluate performance quality from action videos. Most existing methods follow a ``one-by-one'' paradigm, training a separate mod…
GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
Hao Li, Han Fang, Zixin Pan +8
GeoAnchor introduces a framework that breaks down 3D spatial information from 2D images into position, direction, and geometry latent components, enabling more accurate and interpr…
Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World
Yuzhi Huang, Kairun Wen, Rongxin Gao +14
Humans inhabit a physical 4D world where geometric structure and semantic content evolve over time, constituting a dynamic 4D reality (spatial with temporal dimension). While curre…
WAT: Online Video Understanding Needs Watching Before Thinking
Zifan Han, Hongbo Sun, Jinglin Xu +6
Multimodal Large Language Models (MLLMs) have shown strong capabilities in image understanding, motivating recent efforts to extend them to video reasoning. However, existing Video…
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
Tianyi Gao, Hao Li, Han Fang +8
Referring Expression Comprehension (REC) is a vision-language task that localizes a specific image region based on a textual description. Existing REC benchmarks primarily evaluate…