3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.SD2024
A Toolchain for Comprehensive Audio/Video Analysis Using Deep Learning Based Multimodal Approach (A use case of riot or violent context detection)
Lam Pham, Phat Lam, Tin Nguyen +2
In this paper, we present a toolchain for a comprehensive audio/video analysis by leveraging deep learning based multimodal approach. To this end, different specific tasks of Speec…
eess.AS2024
Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders
Phat Lam, Lam Pham, Truong Nguyen +5
Existing speaker diarization systems typically rely on large amounts of manually annotated data, which is labor-intensive and difficult to obtain, especially in real-world scenario…
cs.SD2024★ 3 cited
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
Lam Pham, Phat Lam, Truong Nguyen +2
In this paper, we propose a deep learning based system for the task of deepfake audio detection. In particular, the draw input audio is first transformed into various spectrograms…