9 papers
OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents
Akashah Shabbir, Muhammad Umer Sheikh, Muhammad Akhtar Munir +8
Recent progress in multimodal reasoning has enabled agents that interpret imagery, connect it with language, and execute structured analytical tasks. Extending these capabilities t…
SB-BEVFusion: Enhancing the Robustness against Sensor Malfunction and Corruptions
Markus Essl, Marta Moscati, Mubashir Noman +4
Multimodal sensor fusion has demonstrated remarkable performance improvements over unimodal approaches in 3D object detection for autonomous vehicles. Typically, existing methods t…
AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
Umair Nawaz, Muhammad Zaigham Zaheer, Ufaq Khan +4
Crops, fisheries and livestock form the backbone of global food production, essential to feed the ever-growing global population. However, these sectors face considerable challenge…
Face-Voice Association with Inductive Bias for Maximum Class Separation
Marta Moscati, Oleksandr Kats, Mubashir Noman +4
Face-voice association is widely studied in multimodal learning and is approached representing faces and voices with embeddings that are close for a same person and well separated…
Linking Faces and Voices Across Languages: Insights from the FAME 2026 Challenge
Marta Moscati, Ahmed Abdullah, Muhammad Saad Saeed +7
Over half of the world's population is bilingual and people often communicate under multilingual scenarios. The Face-Voice Association in Multilingual Environments (FAME) 2026 Chal…
RobustA: Robust Anomaly Detection in Multimodal Data
Salem AlMarri, Muhammad Irzam Liaqat, Muhammad Zaigham Zaheer +3
In recent years, multimodal anomaly detection methods have demonstrated remarkable performance improvements over video-only models. However, real-world multimodal data is often cor…