Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting
arXiv:2309.15649 · doi:10.1109/ASRU57964.2023.10389673
Abstract
We explore the ability of large language models (LLMs) to act as speech recognition post-processors that perform rescoring and error correction. Our first focus is on instruction prompting to let LLMs perform these task without fine-tuning, for which we evaluate different prompting schemes, both zero- and few-shot in-context learning, and a novel task activation prompting method that combines causal instructions and demonstration to increase its context windows. Next, we show that rescoring only by in-context learning with frozen LLMs achieves results that are competitive with rescoring by domain-tuned LMs, using a pretrained first-pass recognition system and rescoring output on two out-of-domain tasks (ATIS and WSJ). By combining prompting techniques with fine-tuning we achieve error rates below the N-best oracle level, showcasing the generalization power of the LLMs.
Accepted to IEEE Automatic Speech Recognition and Understanding (ASRU) 2023. 8 pages. 2nd version revised from Sep 29th's version
References in corpus (5)
- Training language models to follow instructions with human feedback
- Large Language Models are Zero-Shot Reasoners
- Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
- N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space
- HyPoradise: An Open Baseline for Generative Speech Recognition with Large Language Models
Cited by in corpus (5)
- Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
- FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
- HandProxy: Expanding the Affordances of Speech Interfaces in Immersive Environments with a Virtual Proxy Hand
- Boot-and-Feedback Framework for Generalist-Expert Model Collaboration in Breast Ultrasound Diagnosis
- Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection