activity
20232025
most citedLarge Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition

10 citations · 26 across the 9 of their papers we have counts for

collaborators
Showing 2024Show all

6 papers · 1 filter

cs.CL2024★ 10 cited

Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition

Chao-Han Huck Yang, Taejin Park, Yuan Gong +18

Given recent advances in generative AI technology, a key question is how large language models (LLMs) can enhance acoustic modeling tasks using text decoding results from a frozen,…

cs.CL2024★ 1 cited

Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization

Yuchen Hu, Chen Chen, Siyin Wang +2

In this paper, we propose reverse inference optimization (RIO), a simple and effective method designed to enhance the robustness of autoregressive-model-based zero-shot text-to-spe…

cs.CL2024★ 1 cited

Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Chen Chen, Yuchen Hu, Wen Wu +3

In recent years, text-to-speech (TTS) technology has witnessed impressive advancements, particularly with large-scale training datasets, showcasing human-level speech quality and i…

cs.CL2024★ 3 cited

Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models

Yuchen Hu, Chen Chen, Chao-Han Huck Yang +4

We propose an unsupervised adaptation framework, Self-TAught Recognizer (STAR), which leverages unlabeled data to enhance the robustness of automatic speech recognition (ASR) syste…

cs.CL2024

Bayesian Example Selection Improves In-Context Learning for Speech, Text, and Visual Modalities

Siyin Wang, Chao-Han Huck Yang, Ji Wu +1

Large language models (LLMs) can adapt to new tasks through in-context learning (ICL) based on a few examples presented in dialogue history without any model parameter update. Desp…

cs.CL2024★ 3 cited

Large Language Models are Efficient Learners of Noise-Robust Speech Recognition

Yuchen Hu, Chen Chen, Chao-Han Huck Yang +4

Recent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which leverages the rich linguistic knowledg…