3 papers
cs.CV2026
K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos
Khush Attarde, Yusuf Ali, Megha Thukral +3
MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text. A crucial, yet underexplored, application of these models lies i…
cs.RO2026
EVE: A Generator-Verifier System for Generative Policies
Yusuf Ali, Gryphon Patlin, Karthik Kothuri +4
Visuomotor policies based on generative such as diffusion and flow-matching have shown strong performance for robotics applications but degrade under distribution shifts, demonstra…
cs.CV2025
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
Karmesh Yadav, Yusuf Ali, Gunshi Gupta +2
Large vision-language models have recently demonstrated impressive performance in planning and control tasks, driving interest in their application to real-world robotics. However,…