1.4k citations · 1.5k across the 41 of their papers we have counts for
5 papers · 1 filter
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
Haibin Wu, Huang-Cheng Chou, Kai-Wei Chang +5
Speech emotion recognition (SER) is a pivotal technology for human-computer interaction systems. However, 80.77% of SER papers yield results that cannot be reproduced. We develop E…
Towards audio language modeling -- an overview
Haibin Wu, Xuanjun Chen, Yi-Cheng Lin +4
Neural audio codecs are initially introduced to compress audio data into compact codes to reduce transmission latency. Researchers recently discovered the potential of codecs as su…
Prompting and Adapter Tuning for Self-supervised Encoder-Decoder Speech Model
Kai-Wei Chang, Ming-Hsin Chen, Yun-Ping Lin +5
Prompting and adapter tuning have emerged as efficient alternatives to fine-tuning (FT) methods. However, existing studies on speech prompting focused on classification tasks and f…
SpeechPrompt v2: Prompt Tuning for Speech Classification Tasks
Kai-Wei Chang, Yu-Kai Wang, Hua Shen +4
Prompt tuning is a technology that tunes a small set of parameters to steer a pre-trained language model (LM) to directly generate the output for downstream tasks. Recently, prompt…
Ensemble knowledge distillation of self-supervised speech models
Kuan-Po Huang, Tzu-hsun Feng, Yu-Kuan Fu +5
Distilled self-supervised models have shown competitive performance and efficiency in recent years. However, there is a lack of experience in jointly distilling multiple self-super…