collaborators

6 papers

cs.CV2026

Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

Tianyi Gao, Han Fang, Tianyi Ding +9

Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing…

cs.CV2026

GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding

Hao Li, Han Fang, Zixin Pan +8

GeoAnchor introduces a framework that breaks down 3D spatial information from 2D images into position, direction, and geometry latent components, enabling more accurate and interpr…

cs.CV2025

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension

Tianyi Gao, Hao Li, Han Fang +8

Referring Expression Comprehension (REC) is a vision-language task that localizes a specific image region based on a textual description. Existing REC benchmarks primarily evaluate…

cs.CV2025

Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval

Haojian Huang, Kaijing Ma, Jin Chen +8

In the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pr…

cs.CV2025

Structure-Aware Prototype Guided Trusted Multi-View Classification

Haojian Huang, Jiahao Shi, Zhe Liu +4

Trustworthy multi-view classification (TMVC) addresses the challenge of achieving reliable decision-making in complex scenarios where multi-source information is heterogeneous, inc…

cs.CV2025

Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment

Jin Chen, Kaijing Ma, Haojian Huang +4

The development of multi-modal models has been rapidly advancing, with some demonstrating remarkable capabilities. However, annotating video-text pairs remains expensive and insuff…