5 papers
GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
Yikun Wang, Zuyan Liu, Ziyi Wang +3
Current research on agentic visual reasoning enables deep multimodal understanding but primarily focuses on image manipulation tools, leaving a gap toward more general-purpose agen…
EnvX: Agentize Everything with Agentic AI
Linyao Chen, Zimian Peng, Yingxuan Yang +4
The widespread availability of open-source repositories has led to a vast collection of reusable software components, yet their utilization remains manual, error-prone, and disconn…
MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
Dianyi Wang, Siyuan Wang, Zejun Li +6
Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across multi-modal tasks by scaling model size and training data. However, these dense LVLMs incur sig…
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
Dianyi Wang, Wei Song, Yikun Wang +4
Typical large vision-language models (LVLMs) apply autoregressive supervision solely to textual sequences, without fully incorporating the visual modality into the learning process…
VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
Yikun Wang, Siyin Wang, Qinyuan Cheng +5
Recent advancements in Large Vision-Language Models have showcased remarkable capabilities. However, they often falter when confronted with complex reasoning tasks that humans typi…