1 citations · 1 across the 13 of their papers we have counts for
26 papers
Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs
Nithin Rao Koluguri, Sasha Meister, Nikolay Karpov +4
Popular ASR test sets adopt inconsistent conventions for numbers, disfluencies, entities, and casing, while standard normalizers erase the format distinctions users care about. Cur…
Hierarchical Policy Optimization for Simultaneous Translation of Unbounded Speech
Siqi Ouyang, Shuoyang Ding, Oleksii Hrinchuk +4
Simultaneous speech translation (SST) generates translations while receiving partial speech input. Recent advances show that large language models (LLMs) can substantially improve…
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
Andrei Andrusenko, Vladimir Bataev, Lilit Grigoryan +3
Unification of automatic speech recognition (ASR) systems reduces development and maintenance costs, but training a single model to perform well in both offline and low-latency str…
Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
Monica Sekoyan, Nithin Rao Koluguri, Nune Tadevosyan +5
This report introduces Canary-1B-v2, a fast, robust multilingual model for Automatic Speech Recognition (ASR) and Speech-to-Text Translation (AST). Built with a FastConformer encod…
A Study on Regularization-Based Continual Learning Methods for Indic ASR
Gokul Adethya T, S. Jaya Nirmala
Indias linguistic diversity poses significant challenges for developing inclusive Automatic Speech Recognition (ASR) systems. Traditional multilingual models, which require simulta…
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
Raymond Grossman, Taejin Park, Kunal Dhawan +6
We introduce SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while m…