audio-visual fusion 1missing modalities 1multilingual speakers 1multimodal learning 1robustness 1speaker identification 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CV2026
Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge
Marta Moscati, Muhammad Saad Saeed, Marina Zanoni +9
The paper describes the POLY-SIM 2026 challenge, which focuses on developing multimodal speaker identification systems that remain robust when audio or visual data are missing and…
cs.SD2026
TARNet: A Temporal-Aware Multi-Scale Architecture for Closed-Set Speaker Identification
Yassin Terraf, Youssef Iraqi
Closed-Set speaker identification aims to assign a speech utterance to one of a predefined set of enrolled speakers and requires robust modeling of speaker-specific characteristics…