most citedContrastive Speech Mixup for Low-resource Keyword Spotting

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS2024

Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis

Kun Zhou, Shengkui Zhao, Yukun Ma +7

Recent language model-based text-to-speech (TTS) frameworks demonstrate scalability and in-context learning capabilities. However, they suffer from robustness issues due to the acc…

cs.SD2023

Are Soft Prompts Good Zero-shot Learners for Speech Recognition?

Dianwen Ng, Chong Zhang, Ruixi Zhang +7

Large self-supervised pre-trained speech models require computationally expensive fine-tuning for downstream tasks. Soft prompt tuning offers a simple parameter-efficient alternati…

cs.SD2023

ACA-Net: Towards Lightweight Speaker Verification using Asymmetric Cross Attention

Jia Qi Yip, Tuan Truong, Dianwen Ng +7

In this paper, we propose ACA-Net, a lightweight, global context-aware speaker embedding extractor for Speaker Verification (SV) that improves upon existing work by using Asymmetri…

cs.SD20231 cited

Contrastive Speech Mixup for Low-resource Keyword Spotting

Dianwen Ng, Ruixi Zhang, Jia Qi Yip +6

Most of the existing neural-based models for keyword spotting (KWS) in smart devices require thousands of training samples to learn a decent audio representation. However, with the…

cs.SD2023

deHuBERT: Disentangling Noise in a Self-supervised Model for Robust Speech Recognition

Dianwen Ng, Ruixi Zhang, Jia Qi Yip +7

Existing self-supervised pre-trained speech models have offered an effective way to leverage massive unannotated corpora to build good automatic speech recognition (ASR). However,…