11 papers
Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning
Jiahao Shao, Yuanbo Yang, Yiyi Liao +3
Tool-augmented vision-language models increasingly "think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind te…
Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis
Xianhao Chen, Jiarui Hu, Yuanbo Yang +5
Open-vocabulary 3D scene understanding aims to segment 3D scenes beyond predefined categories by transferring semantic knowledge from vision-language models. Existing methods have…
Gen3R: 3D Scene Generation Meets Feed-Forward Reconstruction
Jiaxin Huang, Yuanbo Yang, Bangbang Yang +3
We present Gen3R, a method that bridges the strong priors of foundational reconstruction models and video diffusion models for scene-level 3D generation. We repurpose the VGGT reco…
Denoise to Track: Harnessing Video Diffusion Priors for Robust Correspondence
Tianyu Yuan, Yuanbo Yang, Lin-Zhuo Chen +2
In this work, we introduce HeFT (Head-Frequency Tracker), a zero-shot point tracking framework that leverages the visual priors of pretrained video diffusion models. To better unde…
The Constant Eye: Benchmarking and Bridging Appearance Robustness in Autonomous Driving
Jiabao Wang, Hongyu Zhou, Yuanbo Yang +2
Despite rapid progress, autonomous driving algorithms remain notoriously fragile under Out-of-Distribution (OOD) conditions. We identify a critical decoupling failure in current re…
ReRoPE: Repurposing RoPE for Relative Camera Control
Chunyang Li, Yuanbo Yang, Jiahao Shao +3
Video generation with controllable camera viewpoints is essential for applications such as interactive content creation, gaming, and simulation. Existing methods typically adapt pr…