activity
20172026
most citedSeeing wake words: Audio-visual Keyword Spotting

22 citations · 37 across the 27 of their papers we have counts for

collaborators
Showing eess.ASShow all

21 papers · 1 filter

eess.AS2026

SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification

Junyi Peng, Oldřich Plchot, Xiao Song +9

Modern speaker verification (SV) systems rely on speaker embeddings that are effective but difficult to interpret or query in natural language. Most existing speech-text corpora ta…

eess.AS2026

Evaluating voice anonymisation using similarity rank disclosure

Shilpa Chandra, Matteo Pettenò, Nicholas Evans +7

The evaluation of voice anonymisation remains challenging. Current practice relies on automatic speaker verification metrics such as the equal error rate (EER). Performance estimat…

eess.AS2025

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

Sonal Kumar, Šimon Sedláček, Vaibhavi Lokegaonkar +31

Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio unde…

eess.AS2025

Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing

Junyi Peng, Lin Zhang, Jiangyu Han +5

Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on…

eess.AS2025

Analysis of ABC Frontend Audio Systems for the NIST-SRE24

Sara Barahona, Anna Silnova, Ladislav Mošner +14

We present a comprehensive analysis of the embedding extractors (frontends) developed by the ABC team for the audio track of NIST SRE 2024. We follow the two scenarios imposed by N…

eess.AS2024

State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data

Sara Barahona, Ladislav Mošner, Themos Stafylakis +4

In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source Vox…