collaborators

5 papers

cs.RO2026

DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?

Jadelynn Dao, Milan Ganai, Yasmina Abukhadra +7

Vision-Language Models (VLMs) are increasingly deployed as high-level planners for embodied agents, with an emerging strategy of scaling test-time compute to improve capability. Ho…

cs.CV2026

VISTAQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence

Mozhgan Nasr Azadani, Yimu Wang, Yongpeng Zhu +5

Establishing a clear link between model predictions and the visual evidence that supports them is critical for transparency and reliability in multimodal reasoning, yet current mul…

cs.RO2026

Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning

Milan Ganai, Katie Luo, Jonas Frey +2

Embodied Chain-of-Thought (CoT) reasoning has significantly enhanced Vision-Language-Action (VLA) models, yet current methods rely on rigid templates to specify reasoning primitive…

cs.RO2026

Foundation models on the bridge: Semantic hazard detection and safety maneuvers for maritime autonomy with vision-language models

Kim Alexander Christensen, Andreas Gudahl Tufte, Alexey Gusev +5

The draft IMO MASS Code requires autonomous and remotely supervised maritime vessels to detect departures from their operational design domain, enter a predefined fallback that not…

cs.RO2025

Real-Time Out-of-Distribution Failure Prevention via Multi-Modal Reasoning

Milan Ganai, Rohan Sinha, Christopher Agia +3

While foundation models offer promise toward improving robot safety in out-of-distribution (OOD) scenarios, how to effectively harness their generalist knowledge for real-time, dyn…