activity
20242026
collaborators

10 papers

cs.RO2026

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures

Hanan Gani, Tejal Kulkarni, Madhoolika Chodavarapu +2

Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional.…

cs.CV2026

Recovering Cloud Microstructures with Cascaded Diffusion Inversion

Hanan Gani, Guy Pulik, Daniel Rosenfeld +2

High-resolution satellite imagery is critical for observing fine-scale cloud structures that inform weather modification strategies like cloud seeding for rain-enhancement. However…

cs.CV2025

VideoMolmo: Spatio-Temporal Grounding Meets Pointing

Ghazi Shazan Ahmad, Ahmed Heakl, Hanan Gani +4

Spatio-temporal localization is vital for precise interactions across diverse domains, from biological research to autonomous navigation and interactive interfaces. Current video-b…

eess.AS2025

Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs

Sanjoy Chowdhury, Hanan Gani, Nishit Anand +5

Recent advancements in reasoning optimization have greatly enhanced the performance of large language models (LLMs). However, existing work fails to address the complexities of aud…

cs.CV2025

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Shehan Munasinghe, Hanan Gani, Wenqi Zhu +4

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle bas…

cs.CV2025

VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs

Rohit Bharadwaj, Hanan Gani, Muzammal Naseer +2

The recent developments in Large Multi-modal Video Models (Video-LMMs) have significantly enhanced our ability to interpret and analyze video data. Despite their impressive capabil…