7 papers
Multimodal Speaker Verification as a Threat to Speaker Anonymization
Ashi Garg, Cristina Aggazzotti, Leibny Paola GarcÃa-Perera +1
Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulat…
STEB: Style Text Embedding Benchmark
Rafael Rivera Soto, Anna Wegmann, Cristina Aggazzotti
While semantic embeddings are rigorously evaluated on the Massive Text Embedding Benchmark, the evaluation of style embeddings remains fragmented, with each work relying on their o…
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
Content Anonymization for Privacy in Long-form Audio
Cristina Aggazzotti, Ashi Garg, Zexin Cai +1
Voice anonymization techniques have been found to successfully obscure a speaker's acoustic identity in short, isolated utterances in benchmarks such as the VoicePrivacy Challenge.…
The Impact of Automatic Speech Transcription on Speaker Attribution
Cristina Aggazzotti, Matthew Wiesner, Elizabeth Allyn Smith +1
Speaker attribution from speech transcripts is the task of identifying a speaker from the transcript of their speech based on patterns in their language use. This task is especiall…
A stylometric analysis of speaker attribution from speech transcripts
Cristina Aggazzotti, Elizabeth Allyn Smith
Forensic scientists often need to identify an unknown speaker or writer in cases such as ransom calls, covert recordings, alleged suicide notes, or anonymous online communications,…