4 papers
Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning
Jiahao Shao, Yuanbo Yang, Yiyi Liao +3
Tool-augmented vision-language models increasingly "think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind te…
The Constant Eye: Benchmarking and Bridging Appearance Robustness in Autonomous Driving
Jiabao Wang, Hongyu Zhou, Yuanbo Yang +2
Despite rapid progress, autonomous driving algorithms remain notoriously fragile under Out-of-Distribution (OOD) conditions. We identify a critical decoupling failure in current re…
ReRoPE: Repurposing RoPE for Relative Camera Control
Chunyang Li, Yuanbo Yang, Jiahao Shao +3
Video generation with controllable camera viewpoints is essential for applications such as interactive content creation, gaming, and simulation. Existing methods typically adapt pr…
Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation
Zhen Xu, Hongyu Zhou, Sida Peng +13
Depth estimation is a fundamental task in 3D computer vision, crucial for applications such as 3D reconstruction, free-viewpoint rendering, robotics, autonomous driving, and AR/VR…