6 papers
Speak or Stay Silent: Context-Aware Turn-Taking in Multi-Party Dialogue
Kratika Bhagtani, Mrinal Anand, Yu Chen Xu +1
Existing voice AI assistants treat every detected pause as an invitation to speak. This works in dyadic dialogue, but in multi-party settings, where an AI assistant participates al…
HISPASpoof: A New Dataset For Spanish Speech Forensics
Maria Risques, Kratika Bhagtani, Amit Kumar Singh Yadav +1
Zero-shot Voice Cloning (VC) and Text-to-Speech (TTS) methods have advanced rapidly, enabling the generation of highly realistic synthetic speech and raising serious concerns about…
Comparative Analysis of ASR Methods for Speech Deepfake Detection
Davide Salvi, Amit Kumar Singh Yadav, Kratika Bhagtani +3
Recent techniques for speech deepfake detection often rely on pre-trained self-supervised models. These systems, initially developed for Automatic Speech Recognition (ASR), have pr…
DiffSSD: A Diffusion-Based Dataset For Speech Forensics
Kratika Bhagtani, Amit Kumar Singh Yadav, Paolo Bestagini +1
Diffusion-based speech generators are ubiquitous. These methods can generate very high quality synthetic speech and several recent incidents report their malicious use. To counter…
FairSSD: Understanding Bias in Synthetic Speech Detectors
Amit Kumar Singh Yadav, Kratika Bhagtani, Davide Salvi +2
Methods that can generate synthetic speech which is perceptually indistinguishable from speech recorded by a human speaker, are easily available. Several incidents report misuse of…
Compression Robust Synthetic Speech Detection Using Patched Spectrogram Transformer
Amit Kumar Singh Yadav, Ziyue Xiang, Kratika Bhagtani +3
Many deep learning synthetic speech generation tools are readily available. The use of synthetic speech has caused financial fraud, impersonation of people, and misinformation to s…