4 papers
GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
Hao Li, Han Fang, Zixin Pan +8
Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images remains a critical challenge. Existing me…
Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
Tianyi Gao, Han Fang, Tianyi Ding +9
Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing…
ProtRLSearch: A Multi-Round Multimodal Protein Search Agent with Large Language Models Trained via Reinforcement Learning
Congying Liu, Taihao Li, Ming Huang +5
Protein analysis tasks arising in healthcare settings often require accurate reasoning under protein sequence constraints, involving tasks such as functional interpretation of dise…
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
Canhui Tang, Zifan Han, Hongbo Sun +7
Multimodal Large Language Models (MLLMs) have demonstrated significant progress in vision-language tasks, yet they still face challenges when processing long-duration video inputs.…