activity
20222024
most citedAdaptive Knowledge Distillation between Text and Speech Pre-trained Models

1 citations · 3 across the 11 of their papers we have counts for

collaborators

11 papers

cs.SD2024

Towards Audio Codec-based Speech Separation

Jia Qi Yip, Shengkui Zhao, Dianwen Ng +2

Recent improvements in neural audio codec (NAC) models have generated interest in adopting pre-trained codecs for a variety of speech processing applications to take advantage of t…

eess.AS2024

Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis

Kun Zhou, Shengkui Zhao, Yukun Ma +7

Recent language model-based text-to-speech (TTS) frameworks demonstrate scalability and in-context learning capabilities. However, they suffer from robustness issues due to the acc…

cs.SD2024

Dataset-Distillation Generative Model for Speech Emotion Recognition

Fabian Ritter-Gutierrez, Kuan-Po Huang, Jeremy H. M Wong +4

Deep learning models for speech rely on large datasets, presenting computational challenges. Yet, performance hinges on training data size. Dataset Distillation (DD) aims to learn…

cs.SD2023

Are Soft Prompts Good Zero-shot Learners for Speech Recognition?

Dianwen Ng, Chong Zhang, Ruixi Zhang +7

Large self-supervised pre-trained speech models require computationally expensive fine-tuning for downstream tasks. Soft prompt tuning offers a simple parameter-efficient alternati…

cs.SD2023

Analysis of Speech Separation Performance Degradation on Emotional Speech Mixtures

Jia Qi Yip, Dianwen Ng, Bin Ma +1

Despite recent strides made in Speech Separation, most models are trained on datasets with neutral emotions. Emotional speech has been known to degrade performance of models in a v…

cs.SD2023

ACA-Net: Towards Lightweight Speaker Verification using Asymmetric Cross Attention

Jia Qi Yip, Tuan Truong, Dianwen Ng +7

In this paper, we propose ACA-Net, a lightweight, global context-aware speaker embedding extractor for Speaker Verification (SV) that improves upon existing work by using Asymmetri…