17 papers
Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge
Marta Moscati, Muhammad Saad Saeed, Marina Zanoni +9
The paper describes the POLY-SIM 2026 challenge, which focuses on developing multimodal speaker identification systems that remain robust when audio or visual data are missing and…
SB-BEVFusion: Enhancing the Robustness against Sensor Malfunction and Corruptions
Markus Essl, Marta Moscati, Mubashir Noman +4
Multimodal sensor fusion has demonstrated remarkable performance improvements over unimodal approaches in 3D object detection for autonomous vehicles. Typically, existing methods t…
POLY-SIM: Polyglot Speaker Identification with Missing Modality Grand Challenge 2026 Evaluation Plan
Marta Moscati, Muhammad Saad Saeed, Marina Zanoni +8
Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing. However, in real-w…
From Native Memes to Global Moderation: Cross-Cultural Evaluation of Vision-Language Models for Hateful Meme Detection
Mo Wang, Kaixuan Ren, Pratik Jalan +5
Cultural context profoundly shapes how people interpret online content, yet vision-language models (VLMs) remain predominantly trained through Western or English-centric lenses. Th…
EoCD: Encoder only Remote Sensing Change Detection
Mubashir Noman, Mustansar Fiaz, Hiyam Debary +4
Being a cornerstone of temporal analysis, change detection has been playing a pivotal role in modern earth observation. Existing change detection methods rely on the Siamese encode…
Robust Harmful Meme Detection under Missing Modalities via Shared Representation Learning
Felix Breiteneder, Mohammad Belal, Muhammad Saad Saeed +5
Internet memes are powerful tools for communication, capable of spreading political, psychological, and sociocultural ideas. However, they can be harmful and can be used to dissemi…