collaborators

26 papers

cs.CV2026

K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos

Khush Attarde, Yusuf Ali, Megha Thukral +3

MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text. A crucial, yet underexplored, application of these models lies i…

cs.RO2026

SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation

Chengyue Huang, Khang Vo Huynh, Sebastian Elbaum +2

Robotic manipulation is typically evaluated by task success, but successful completion does not guarantee safe execution. Many safety failures are temporal: a robot may touch a cle…

cs.RO2026

Video2Sim2Real: Full-Stack Autonomous Dexterous Skill Acquisition from a Single Human Video

Yunhai Han, Jianuo Qiu, Linhao Bai +14

Human manipulation videos are a convenient and intuitive source for robot learning. However, directly transferring human dexterity to robots remains challenging due to perception e…

cs.RO2026

EVE: A Generator-Verifier System for Generative Policies

Yusuf Ali, Gryphon Patlin, Karthik Kothuri +4

Visuomotor policies based on generative such as diffusion and flow-matching have shown strong performance for robotics applications but degrade under distribution shifts, demonstra…

cs.RO2026

SeeTraceAct: Visibility-Aware Latent Planning from Cross-Embodiment Demonstration Videos

Jaehyeon Son, Junhyun Kim, Kyle Kam +7

Vision-language-action models (VLAs) are promising general-purpose robot policies, but adapting them to new tasks typically requires costly task-specific teleoperation data. As an…

cs.RO2026

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring

Seongheon Park, Wendi Li, Changdae Oh +4

Vision-Language-Action (VLA) models enable robots to follow natural language instructions and generalize across diverse tasks, but they remain vulnerable to execution failures that…