activity
20162023
most citedTime-domain speaker extraction network

36 citations · 123 across the 43 of their papers we have counts for

collaborators
Showing 2022Show all

12 papers · 1 filter

cs.CL2022

Punctuation Restoration for Singaporean Spoken Languages: English, Malay, and Mandarin

Abhinav Rao, Ho Thi-Nga, Chng Eng-Siong

This paper presents the work of restoring punctuation for ASR transcripts generated by multilingual ASR systems. The focus languages are English, Mandarin, and Malay which are thre…

cs.SD2022

I2CR: Improving Noise Robustness on Keyword Spotting Using Inter-Intra Contrastive Regularization

Dianwen Ng, Jia Qi Yip, Tanmay Surana +6

Noise robustness in keyword spotting remains a challenge as many models fail to overcome the heavy influence of noises, causing the deterioration of the quality of feature embeddin…

eess.AS2022

DENT-DDSP: Data-efficient noisy speech generator using differentiable digital signal processors for explicit distortion modelling and noise-robust speech recognition

Z. Guo, C. Chen, E. S. Chng

The performances of automatic speech recognition (ASR) systems degrade drastically under noisy conditions. Explicit distortion modelling (EDM), as a feature compensation step, is a…

cs.SD2022★ 5 cited

Continual Learning For On-Device Environmental Sound Classification

Yang Xiao, Xubo Liu, James King +4

Continuously learning new classes without catastrophic forgetting is a challenging problem for on-device environmental sound classification given the restrictions on computation re…

cs.SD2022

Language-Based Audio Retrieval with Converging Tied Layers and Contrastive Loss

Andrew Koh, Eng Siong Chng

In this paper, we tackle the new Language-Based Audio Retrieval task proposed in DCASE 2022. Firstly, we introduce a simple, scalable architecture which ties both the audio and tex…

cs.CL2022

Automated Audio Captioning with Epochal Difficult Captions for Curriculum Learning

Andrew Koh, Soham Tiwari, Chng Eng Siong

In this paper, we propose an algorithm, Epochal Difficult Captions, to supplement the training of any model for the Automated Audio Captioning task. Epochal Difficult Captions is a…