14 citations · 41 across the 23 of their papers we have counts for
5 papers · 1 filter
LibriSpeech-PC: Benchmark for Evaluation of Punctuation and Capitalization Capabilities of end-to-end ASR Models
Aleksandr Meister, Matvei Novikov, Nikolay Karpov +3
Traditional automatic speech recognition (ASR) models output lower-cased words without punctuation marks, which reduces readability and necessitates a subsequent text processing mo…
Retrieval meets Long Context Large Language Models
Peng Xu, Wei Ping, Xianchao Wu +7
Extending the context window of large language models (LLMs) is getting popular recently, while the solution of augmenting LLMs with retrieval has existed for years. The natural qu…
A Chat About Boring Problems: Studying GPT-based text normalization
Yang Zhang, Travis M. Bartley, Mariana Graterol-Fuenmayor +3
Text normalization - the conversion of text from written to spoken form - is traditionally assumed to be an ill-formed task for language models. In this work, we argue otherwise. W…
SpellMapper: A non-autoregressive neural spellchecker for ASR customization with candidate retrieval based on n-gram mappings
Alexandra Antonova, Evelina Bakhturina, Boris Ginsburg
Contextual spelling correction models are an alternative to shallow fusion to improve automatic speech recognition (ASR) quality given user vocabulary. To deal with large user voca…
Automatic Heteronym Resolution Pipeline Using RAD-TTS Aligners
Jocelyn Huang, Evelina Bakhturina, Oktai Tatanov
Grapheme-to-phoneme (G2P) transduction is part of the standard text-to-speech (TTS) pipeline. However, G2P conversion is difficult for languages that contain heteronyms -- words th…