activity
20222025
most citedContinual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems

4 citations · 5 across the 10 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2025

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission

Nirmalya Mallick Thakur, Jia Qi Yip, Eng Siong Chng

Neural audio codecs (NACs) have made significant advancements in recent years and are rapidly being adopted in many audio processing pipelines. However, they can introduce audio di…

cs.SD2024

Towards Audio Codec-based Speech Separation

Jia Qi Yip, Shengkui Zhao, Dianwen Ng +2

Recent improvements in neural audio codec (NAC) models have generated interest in adopting pre-trained codecs for a variety of speech processing applications to take advantage of t…

cs.SD2023

Analysis of Speech Separation Performance Degradation on Emotional Speech Mixtures

Jia Qi Yip, Dianwen Ng, Bin Ma +1

Despite recent strides made in Speech Separation, most models are trained on datasets with neutral emotions. Emotional speech has been known to degrade performance of models in a v…

cs.SD2023

ACA-Net: Towards Lightweight Speaker Verification using Asymmetric Cross Attention

Jia Qi Yip, Tuan Truong, Dianwen Ng +7

In this paper, we propose ACA-Net, a lightweight, global context-aware speaker embedding extractor for Speaker Verification (SV) that improves upon existing work by using Asymmetri…

cs.SD20231 cited

Contrastive Speech Mixup for Low-resource Keyword Spotting

Dianwen Ng, Ruixi Zhang, Jia Qi Yip +6

Most of the existing neural-based models for keyword spotting (KWS) in smart devices require thousands of training samples to learn a decent audio representation. However, with the…

cs.SD2023

deHuBERT: Disentangling Noise in a Self-supervised Model for Robust Speech Recognition

Dianwen Ng, Ruixi Zhang, Jia Qi Yip +7

Existing self-supervised pre-trained speech models have offered an effective way to leverage massive unannotated corpora to build good automatic speech recognition (ASR). However,…