works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CV2026

Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

Tianyi Gao, Han Fang, Tianyi Ding +9

Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing…

cs.CV2026

GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding

Hao Li, Han Fang, Zixin Pan +8

GeoAnchor introduces a framework that breaks down 3D spatial information from 2D images into position, direction, and geometry latent components, enabling more accurate and interpr…

cs.AI2026

Segment-Aligned Policy Optimization for Multi-Modal Reasoning

Lei Gao, Zhuoming Li, Mengxi Jia +4

Existing reinforcement learning approaches for Large Language Models typically perform policy optimization at the granularity of individual tokens or entire response sequences. How…

cs.CV2025

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension

Tianyi Gao, Hao Li, Han Fang +8

Referring Expression Comprehension (REC) is a vision-language task that localizes a specific image region based on a textual description. Existing REC benchmarks primarily evaluate…

cs.CV2025

TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding

Canhui Tang, Zifan Han, Hongbo Sun +7

Multimodal Large Language Models (MLLMs) have demonstrated significant progress in vision-language tasks, yet they still face challenges when processing long-duration video inputs.…

cs.CV2025

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

Peng Xu, Shengwu Xiong, Jiajun Zhang +125

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…