3 papers
eess.AS2025
State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data
Sara Barahona, Ladislav Mošner, Themos Stafylakis +4
In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source Vox…
eess.AS2025
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
Sonal Kumar, Å imon SedláÄek, Vaibhavi Lokegaonkar +31
Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio unde…
eess.AS2025
Analysis of ABC Frontend Audio Systems for the NIST-SRE24
Sara Barahona, Anna Silnova, Ladislav Mošner +14
We present a comprehensive analysis of the embedding extractors (frontends) developed by the ABC team for the audio track of NIST SRE 2024. We follow the two scenarios imposed by N…