26 papers
K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos
Khush Attarde, Yusuf Ali, Megha Thukral +3
MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text. A crucial, yet underexplored, application of these models lies i…
SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation
Chengyue Huang, Khang Vo Huynh, Sebastian Elbaum +2
Robotic manipulation is typically evaluated by task success, but successful completion does not guarantee safe execution. Many safety failures are temporal: a robot may touch a cle…
Video2Sim2Real: Full-Stack Autonomous Dexterous Skill Acquisition from a Single Human Video
Yunhai Han, Jianuo Qiu, Linhao Bai +14
Human manipulation videos are a convenient and intuitive source for robot learning. However, directly transferring human dexterity to robots remains challenging due to perception e…
EVE: A Generator-Verifier System for Generative Policies
Yusuf Ali, Gryphon Patlin, Karthik Kothuri +4
Visuomotor policies based on generative such as diffusion and flow-matching have shown strong performance for robotics applications but degrade under distribution shifts, demonstra…
SeeTraceAct: Visibility-Aware Latent Planning from Cross-Embodiment Demonstration Videos
Jaehyeon Son, Junhyun Kim, Kyle Kam +7
Vision-language-action models (VLAs) are promising general-purpose robot policies, but adapting them to new tasks typically requires costly task-specific teleoperation data. As an…
Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring
Seongheon Park, Wendi Li, Changdae Oh +4
Vision-Language-Action (VLA) models enable robots to follow natural language instructions and generalize across diverse tasks, but they remain vulnerable to execution failures that…