From the 1 of 12 linked papers with an AI index.
12 papers
Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge
Marta Moscati, Muhammad Saad Saeed, Marina Zanoni +9
The paper describes the POLY-SIM 2026 challenge, which focuses on developing multimodal speaker identification systems that remain robust when audio or visual data are missing and…
SB-BEVFusion: Enhancing the Robustness against Sensor Malfunction and Corruptions
Markus Essl, Marta Moscati, Mubashir Noman +4
Multimodal sensor fusion has demonstrated remarkable performance improvements over unimodal approaches in 3D object detection for autonomous vehicles. Typically, existing methods t…
Adaptive Autoguidance for Item-Side Fairness in Diffusion Recommender Systems
Zihan Li, Gustavo Escobedo, Marta Moscati +2
Diffusion recommender systems achieve strong recommendation accuracy but often suffer from popularity bias, resulting in unequal item exposure. To address this shortcoming, we intr…
POLY-SIM: Polyglot Speaker Identification with Missing Modality Grand Challenge 2026 Evaluation Plan
Marta Moscati, Muhammad Saad Saeed, Marina Zanoni +8
Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing. However, in real-w…
Face-Voice Association with Inductive Bias for Maximum Class Separation
Marta Moscati, Oleksandr Kats, Mubashir Noman +4
Face-voice association is widely studied in multimodal learning and is approached representing faces and voices with embeddings that are close for a same person and well separated…
Linking Faces and Voices Across Languages: Insights from the FAME 2026 Challenge
Marta Moscati, Ahmed Abdullah, Muhammad Saad Saeed +7
Over half of the world's population is bilingual and people often communicate under multilingual scenarios. The Face-Voice Association in Multilingual Environments (FAME) 2026 Chal…