4 papers
Environmental Sound Deepfake Detection Using Deep-Learning Framework
Khoi Vu, Dat Tran, Khanh Do +8
In this paper, we propose a deep-learning framework for Environmental Sound Deepfake Detection (ESDD) - the task of identifying whether the sound scene and sound event in an input…
A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
Lam Pham, Phat Lam, Dat Tran +6
Thanks to advancements in deep learning, speech generation systems now power a variety of real-world applications, such as text-to-speech for individuals with speech disorders, voi…
A Toolchain for Comprehensive Audio/Video Analysis Using Deep Learning Based Multimodal Approach (A use case of riot or violent context detection)
Lam Pham, Phat Lam, Tin Nguyen +2
In this paper, we present a toolchain for a comprehensive audio/video analysis by leveraging deep learning based multimodal approach. To this end, different specific tasks of Speec…
Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders
Phat Lam, Lam Pham, Truong Nguyen +5
Existing speaker diarization systems typically rely on large amounts of manually annotated data, which is labor-intensive and difficult to obtain, especially in real-world scenario…