activity
20242026
collaborators

12 papers

cs.CV2026

Divide-and-Conquer Approach to Holistic Cognition in High-Similarity Contexts with Limited Data

Shijie Wang, Zijian Wang, Yadan Luo +3

Ultra-fine-grained visual categorization (Ultra-FGVC) aims to classify highly similar subcategories within fine-grained objects using limited training samples. However, holistic ye…

cs.RO2026

AnchorVLA: Anchored Diffusion for Efficient End-to-End Mobile Manipulation

Jia Syuen Lim, Zhizhen Zhang, Peter Bohm +3

A central challenge in mobile manipulation is preserving multiple plausible action models while remaining reactive during execution. A bottle in a cluttered scene can often be appr…

cs.CV2026

VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model

Xiangyu Sun, Shijie Wang, Fengyi Zhang +5

World models that forecast scene evolution by generating future video frames devote the bulk of their capacity to photometric details, yet the resulting predictions often remain ge…

cs.CV2025

Language-driven Fine-grained Retrieval

Shijie Wang, Xin Yu, Yadan Luo +3

Existing fine-grained image retrieval (FGIR) methods learn discriminative embeddings by adopting semantically sparse one-hot labels derived from category names as supervision. Whil…

cs.CV2025

TALO: Pushing 3D Vision Foundation Models Towards Globally Consistent Online Reconstruction

Fengyi Zhang, Tianjun Zhang, Kasra Khosoussi +3

3D vision foundation models have shown strong generalization in reconstructing key 3D attributes from uncalibrated images through a single feed-forward pass. However, when deployed…

cs.RO2025

MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent

Yuxia Fu, Zhizhen Zhang, Yuqi Zhang +3

Recent Vision-Language-Action (VLA) models reformulate vision-language models by tuning them with millions of robotic demonstrations. While they perform well when fine-tuned for a…