12 papers
Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge
Marta Moscati, Muhammad Saad Saeed, Marina Zanoni +9
The paper describes the POLY-SIM 2026 challenge, which focuses on developing multimodal speaker identification systems that remain robust when audio or visual data are missing and…
Towards Effective Waste Segmentation for Automated Waste Recycling in Cluttered Background
Mamoona Javaid, Mubashir Noman, Abdul Hannan +3
Rapid expansion of urban areas and population growth is causing an immense increase in waste production, which demands the need for efficient and automated waste management. In thi…
SB-BEVFusion: Enhancing the Robustness against Sensor Malfunction and Corruptions
Markus Essl, Marta Moscati, Mubashir Noman +4
Multimodal sensor fusion has demonstrated remarkable performance improvements over unimodal approaches in 3D object detection for autonomous vehicles. Typically, existing methods t…
POLY-SIM: Polyglot Speaker Identification with Missing Modality Grand Challenge 2026 Evaluation Plan
Marta Moscati, Muhammad Saad Saeed, Marina Zanoni +8
Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing. However, in real-w…
EoCD: Encoder only Remote Sensing Change Detection
Mubashir Noman, Mustansar Fiaz, Hiyam Debary +4
Being a cornerstone of temporal analysis, change detection has been playing a pivotal role in modern earth observation. Existing change detection methods rely on the Siamese encode…
Distillation-based Layer Dropping (DLD): Effective End-to-end Framework for Dynamic Speech Networks
Abdul Hannan, Daniele Falavigna, Shah Nawaz +3
Edge devices operate in constrained and varying resource settings, requiring dynamic architectures that can adapt to limitations of the available resources. To meet such demands, l…