works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CV2026

Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

Tianyi Gao, Han Fang, Tianyi Ding +9

Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing…

cs.CV2026

MoAKE: Toward Unified All-in-One Action Quality Assessment via Mixture of Action Knowledge Experts

Huangbiao Xu, Huanqi Wu, Xiao Ke +3

Action Quality Assessment (AQA) aims to objectively evaluate performance quality from action videos. Most existing methods follow a ``one-by-one'' paradigm, training a separate mod…

cs.CV2026

GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding

Hao Li, Han Fang, Zixin Pan +8

GeoAnchor introduces a framework that breaks down 3D spatial information from 2D images into position, direction, and geometry latent components, enabling more accurate and interpr…

cs.CV2026

Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World

Yuzhi Huang, Kairun Wen, Rongxin Gao +14

Humans inhabit a physical 4D world where geometric structure and semantic content evolve over time, constituting a dynamic 4D reality (spatial with temporal dimension). While curre…

cs.CV2026

WAT: Online Video Understanding Needs Watching Before Thinking

Zifan Han, Hongbo Sun, Jinglin Xu +6

Multimodal Large Language Models (MLLMs) have shown strong capabilities in image understanding, motivating recent efforts to extend them to video reasoning. However, existing Video…

cs.CV2025

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension

Tianyi Gao, Hao Li, Han Fang +8

Referring Expression Comprehension (REC) is a vision-language task that localizes a specific image region based on a textual description. Existing REC benchmarks primarily evaluate…