8 papers
SABRE: Scalable and Automated Benchmarking of VLMs under Stress
Zixuan Lan, Luzhe Sun, Matthew R. Walter +1
Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisf…
Flow-based Policy Adaptation without Policy Updates
Luzhe Sun, Jingtian Ji, Haoran Chen +2
Leveraging prior knowledge from pretrained policies, foundation models, or human operators offers an efficient alternative to learning robot skills from scratch. However, these age…
What Are We Actually Benchmarking in Robot Manipulation?
Tianchong Jiang, Xiangshan Tan, Samuel Wheeler +3
A robotics benchmark score measures success under one fixed evaluation setup, yet is routinely treated as evidence of general manipulation capability. We identify four failure mode…
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?
Zixuan Lan, Luzhe Sun, Matthew R. Walter +1
Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to what extent such scores truly re…
HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token
Sai Akhil Kogilathota, Sripadha Vallabha E G, Luzhe Sun +1
Hallucinations remain a persistent challenge for vision-language models (VLMs), which often describe nonexistent objects or fabricate facts. Existing detection methods typically op…
To the Noise and Back: Diffusion for Shared Autonomy
Takuma Yoneda, Luzhe Sun, Ge Yang +2
Shared autonomy is an operational concept in which a user and an autonomous agent collaboratively control a robotic system. It provides a number of advantages over the extremes of…