2 papers
cs.RO2025
TReF-6: Inferring Task-Relevant Frames from a Single Demonstration for One-Shot Skill Generalization
Yuxuan Ding, Shuangge Wang, Tesca Fitzgerald
Robots often struggle to generalize from a single demonstration due to the lack of a transferable and interpretable spatial representation. In this work, we introduce TReF-6, a met…
cs.CV2024
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Ziyao Shangguan, Chuhan Li, Yuxuan Ding +4
Existing benchmarks often highlight the remarkable performance achieved by state-of-the-art Multimodal Foundation Models (MFMs) in leveraging temporal context for video understandi…