36 citations · 123 across the 43 of their papers we have counts for
12 papers · 1 filter
Punctuation Restoration for Singaporean Spoken Languages: English, Malay, and Mandarin
Abhinav Rao, Ho Thi-Nga, Chng Eng-Siong
This paper presents the work of restoring punctuation for ASR transcripts generated by multilingual ASR systems. The focus languages are English, Mandarin, and Malay which are thre…
I2CR: Improving Noise Robustness on Keyword Spotting Using Inter-Intra Contrastive Regularization
Dianwen Ng, Jia Qi Yip, Tanmay Surana +6
Noise robustness in keyword spotting remains a challenge as many models fail to overcome the heavy influence of noises, causing the deterioration of the quality of feature embeddin…
DENT-DDSP: Data-efficient noisy speech generator using differentiable digital signal processors for explicit distortion modelling and noise-robust speech recognition
Z. Guo, C. Chen, E. S. Chng
The performances of automatic speech recognition (ASR) systems degrade drastically under noisy conditions. Explicit distortion modelling (EDM), as a feature compensation step, is a…
Continual Learning For On-Device Environmental Sound Classification
Yang Xiao, Xubo Liu, James King +4
Continuously learning new classes without catastrophic forgetting is a challenging problem for on-device environmental sound classification given the restrictions on computation re…
Language-Based Audio Retrieval with Converging Tied Layers and Contrastive Loss
Andrew Koh, Eng Siong Chng
In this paper, we tackle the new Language-Based Audio Retrieval task proposed in DCASE 2022. Firstly, we introduce a simple, scalable architecture which ties both the audio and tex…
Automated Audio Captioning with Epochal Difficult Captions for Curriculum Learning
Andrew Koh, Soham Tiwari, Chng Eng Siong
In this paper, we propose an algorithm, Epochal Difficult Captions, to supplement the training of any model for the Automated Audio Captioning task. Epochal Difficult Captions is a…