6 papers
RECALL: Recovery Experience Collection for Active Lifelong Learning in Vision-Language-Action Models
Ulas Berk Karli, Tesca Fitzgerald
Vision-Language-Action (VLA) models are commonly fine-tuned through passive imitation learning, where additional demonstrations are collected for tasks where the policy performs po…
INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models
Ulas Berk Karli, Ziyao Shangguan, Tesca FItzgerald
Recent Vision-Language-Action (VLA) models show strong generalization capabilities, yet they lack introspective mechanisms for anticipating failures and requesting help from a huma…
Enhancing Goal Inference via Correction Timing
Anjiabei Wang, Shuangge Wang, Tesca Fitzgerald
Corrections offer a natural modality for people to provide feedback to a robot, by (i) intervening in the robot's behavior when they believe the robot is failing (or will fail) the…
TReF-6: Inferring Task-Relevant Frames from a Single Demonstration for One-Shot Skill Generalization
Yuxuan Ding, Shuangge Wang, Tesca Fitzgerald
Robots often struggle to generalize from a single demonstration due to the lack of a transferable and interpretable spatial representation. In this work, we introduce TReF-6, a met…
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Ziyao Shangguan, Chuhan Li, Yuxuan Ding +4
Existing benchmarks often highlight the remarkable performance achieved by state-of-the-art Multimodal Foundation Models (MFMs) in leveraging temporal context for video understandi…
Effects of Robot Competency and Motion Legibility on Human Correction Feedback
Shuangge Wang, Anjiabei Wang, Sofiya Goncharova +2
As robot deployments become more commonplace, people are likely to take on the role of supervising robots (i.e., correcting their mistakes) rather than directly teaching them. Prio…