FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
arXiv:2501.12979 · doi:10.1109/SLT61566.2024.10832257
Abstract
In this paper, we present an encoder-decoder model leveraging Flan-T5 for post-Automatic Speech Recognition (ASR) Generative Speech Error Correction (GenSEC), and we refer to it as FlanEC. We explore its application within the GenSEC framework to enhance ASR outputs by mapping n-best hypotheses into a single output sentence. By utilizing n-best lists from ASR models, we aim to improve the linguistic correctness, accuracy, and grammaticality of final ASR transcriptions. Specifically, we investigate whether scaling the training data and incorporating diverse datasets can lead to significant improvements in post-ASR error correction. We evaluate FlanEC using the HyPoradise dataset, providing a comprehensive analysis of the model's effectiveness in this domain. Furthermore, we assess the proposed approach under different settings to evaluate model scalability and efficiency, offering valuable insights into the potential of instruction-tuned encoder-decoder models for this task.
Accepted at the 2024 IEEE Workshop on Spoken Language Technology (SLT) - GenSEC Challenge
References in corpus (13)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- LoRA: Low-Rank Adaptation of Large Language Models
- Scaling Instruction-Finetuned Language Models
- Emergent Abilities of Large Language Models
- Lip Reading Sentences in the Wild
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
- TED-LIUM 3: twice as much data and corpus repartition for experiments on speaker adaptation
- Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting
- N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space
- HyPoradise: An Open Baseline for Generative Speech Recognition with Large Language Models
- Quantifying the Role of Textual Predictability in Automatic Speech Recognition