activity
20242026
most citedTOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models

1 citations · 1 across the 6 of their papers we have counts for

collaborators
Showing cs.ROShow all

5 papers · 1 filter

cs.RO2026

RECALL: Recovery Experience Collection for Active Lifelong Learning in Vision-Language-Action Models

Ulas Berk Karli, Tesca Fitzgerald

Vision-Language-Action (VLA) models are commonly fine-tuned through passive imitation learning, where additional demonstrations are collected for tasks where the policy performs po…

cs.RO2026

Enhancing Goal Inference via Correction Timing

Anjiabei Wang, Shuangge Wang, Tesca Fitzgerald

Corrections offer a natural modality for people to provide feedback to a robot, by (i) intervening in the robot's behavior when they believe the robot is failing (or will fail) the…

cs.RO2025

INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models

Ulas Berk Karli, Ziyao Shangguan, Tesca FItzgerald

Recent Vision-Language-Action (VLA) models show strong generalization capabilities, yet they lack introspective mechanisms for anticipating failures and requesting help from a huma…

cs.RO2025

TReF-6: Inferring Task-Relevant Frames from a Single Demonstration for One-Shot Skill Generalization

Yuxuan Ding, Shuangge Wang, Tesca Fitzgerald

Robots often struggle to generalize from a single demonstration due to the lack of a transferable and interpretable spatial representation. In this work, we introduce TReF-6, a met…

cs.RO2025

Effects of Robot Competency and Motion Legibility on Human Correction Feedback

Shuangge Wang, Anjiabei Wang, Sofiya Goncharova +2

As robot deployments become more commonplace, people are likely to take on the role of supervising robots (i.e., correcting their mistakes) rather than directly teaching them. Prio…