7 papers
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
Matthew Foutter, Matteo Cercola, Lena Wild +4
Embodied Chain-of-Thought has emerged as a promising mechanism to enhance robot decision-making and interpretability in black-box Vision-Language Action (VLA) models. However, whet…
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions
Longfei Li, Zhiwen Fan, Wenyan Cong +10
Synthesizing realistic Martian landscape videos is crucial for mission rehearsal and robotic simulation. However, this task poses unique challenges due to the scarcity of high-qual…
RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
Jacky Kwok, Christopher Agia, Rohan Sinha +5
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in visuomotor control, yet ensuring their robustness in unstructured real-world environments remains a…
Vision Foundation Model Embedding-Based Semantic Anomaly Detection
Max Peter Ronecker, Matthew Foutter, Amine Elhafsi +4
Semantic anomalies are contextually invalid or unusual combinations of familiar visual elements that can cause undefined behavior and failures in system-level reasoning for autonom…
Space-LLaVA: a Vision-Language Model Adapted to Extraterrestrial Applications
Matthew Foutter, Daniele Gammelli, Justin Kruger +5
Foundation Models (FMs), e.g., large language models, possess attributes of intelligence which offer promise to endow a robot with the contextual understanding necessary to navigat…
Realistic Extreme Behavior Generation for Improved AV Testing
Robert Dyro, Matthew Foutter, Ruolin Li +4
This work introduces a framework to diagnose the strengths and shortcomings of Autonomous Vehicle (AV) collision avoidance technology with synthetic yet realistic potential collisi…