10 papers · 1 filter
Observing and Controlling Features in Vision-Language-Action Models
Hugo Buurmeijer, Carmen Amo Alonso, Aiden Swann +1
Vision-Language-Action Models (VLAs) have shown remarkable progress towards embodied intelligence. While their architecture partially resembles that of Large Language Models (LLMs)…
Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment
Jacky Kwok, Xilun Zhang, Mengdi Xu +4
The long-standing vision of general-purpose robots hinges on their ability to understand and act upon natural language instructions. Vision-Language-Action (VLA) models have made r…
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
Jacky Kwok, Christopher Agia, Rohan Sinha +5
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in visuomotor control, yet ensuring their robustness in unstructured real-world environments remains a…
CUPID: Curating Data your Robot Loves with Influence Functions
Christopher Agia, Rohan Sinha, Jingyun Yang +5
In robot imitation learning, policy performance is tightly coupled with the quality and composition of the demonstration data. Yet, developing a precise understanding of how indivi…
Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis
Yuan Gao, Mattia Piccinini, Yuchen Zhang +12
For autonomous vehicles, safe navigation in complex environments depends on handling a broad range of diverse and rare driving scenarios. Simulation- and scenario-based testing hav…
Scan, Materialize, Simulate: A Generalizable Framework for Physically Grounded Robot Planning
Amine Elhafsi, Daniel Morton, Marco Pavone
Autonomous robots must reason about the physical consequences of their actions to operate effectively in unstructured, real-world environments. We present Scan, Materialize, Simula…