5 papers
A Toolchain for Comprehensive Audio/Video Analysis Using Deep Learning Based Multimodal Approach (A use case of riot or violent context detection)
Lam Pham, Phat Lam, Tin Nguyen +2
In this paper, we present a toolchain for a comprehensive audio/video analysis by leveraging deep learning based multimodal approach. To this end, different specific tasks of Speec…
Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders
Phat Lam, Lam Pham, Truong Nguyen +5
Existing speaker diarization systems typically rely on large amounts of manually annotated data, which is labor-intensive and difficult to obtain, especially in real-world scenario…
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
Thang M. Pham, Peijie Chen, Tin Nguyen +3
CLIP-based classifiers rely on the prompt containing a {class name} that is known to the text encoder. Therefore, they perform poorly on new classes or the classes whose names rare…
The Impact of Frequency Bands on Acoustic Anomaly Detection of Machines using Deep Learning Based Model
Tin Nguyen, Lam Pham, Phat Lam +3
In this paper, we propose a deep learning based model for Acoustic Anomaly Detection of Machines, the task for detecting abnormal machines by analysing the machine sound. By conduc…
Towards Conceptualization of "Fair Explanation": Disparate Impacts of anti-Asian Hate Speech Explanations on Content Moderators
Tin Nguyen, Jiannan Xu, Aayushi Roy +2
Recent research at the intersection of AI explainability and fairness has focused on how explanations can improve human-plus-AI task performance as assessed by fairness measures. W…