4 papers
CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
Letian Fu, Justin Yu, Karim El-Refai +13
"Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness as autonomous controllers for embodied ma…
MIRAGE: The Illusion of Visual Understanding
Mohammad Asadi, Jack W. O'Sullivan, Fang Cao +5
Multimodal AI systems have achieved remarkable performance across a broad range of real-world tasks, yet the mechanisms underlying visual-language reasoning remain surprisingly poo…
MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management
Jack W O'Sullivan, Mohammad Asadi, Lennart Elbe +8
Cardiovascular disease remains the leading cause of global mortality, with progress hindered by human interpretation of complex cardiac tests. Current AI vision-language models are…
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models
Chen Wang, Fei Xia, Wenhao Yu +6
Learning to perform manipulation tasks from human videos is a promising approach for teaching robots. However, many manipulation tasks require changing control parameters during ta…