collaborators

6 papers

cs.CV2026

SABRE: Scalable and Automated Benchmarking of VLMs under Stress

Zixuan Lan, Luzhe Sun, Matthew R. Walter +1

Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisf…

cs.CV2026

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

Zixuan Lan, Luzhe Sun, Matthew R. Walter +1

Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to what extent such scores truly re…

cs.RO2025

Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning

Tianchong Jiang, Jingtian Ji, Xiangshan Tan +4

We study view-invariant imitation learning by explicitly conditioning policies on camera extrinsics. Using Plucker embeddings of per-pixel rays, we show that conditioning on extrin…

cs.CV2025

FastMap: Revisiting Structure from Motion through First-Order Optimization

Jiahao Li, Haochen Wang, Muhammad Zubair Irshad +4

We propose FastMap, a new global structure from motion method focused on speed and simplicity. Previous methods like COLMAP and GLOMAP are able to estimate high-precision camera po…

cs.GR2025

SplArt: Articulation Estimation and Part-Level Reconstruction with 3D Gaussian Splatting

Shengjie Lin, Jiading Fang, Muhammad Zubair Irshad +4

Reconstructing articulated objects prevalent in daily environments is crucial for applications in augmented/virtual reality and robotics. However, existing methods face scalability…

cs.RO2025

FlashBack: Consistency Model-Accelerated Shared Autonomy

Luzhe Sun, Jingtian Ji, Xiangshan Tan +1

Shared autonomy is an enabling technology that provides users with control authority over robots that would otherwise be difficult if not impossible to directly control. Yet, stand…