most citedCodec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations

7 citations · 10 across the 12 of their papers we have counts for

collaborators
Showing eess.ASShow all

8 papers · 1 filter

eess.AS20247 cited

Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations

Kunal Dhawan, Nithin Rao Koluguri, Ante Jukić +3

Discrete speech representations have garnered recent attention for their efficacy in training transformer-based models for various speech-related tasks such as automatic speech rec…

eess.AS2024

Instruction Data Generation and Unsupervised Adaptation for Speech Language Models

Vahid Noroozi, Zhehuai Chen, Somshubra Majumdar +3

In this paper, we propose three methods for generating synthetic samples to train and evaluate multimodal large language models capable of processing both text and speech inputs. A…

eess.AS2024

Flexible Multichannel Speech Enhancement for Noise-Robust Frontend

Ante Jukić, Jagadeesh Balam, Boris Ginsburg

This paper proposes a flexible multichannel speech enhancement system with the main goal of improving robustness of automatic speech recognition (ASR) in noisy conditions. The prop…

eess.AS2023

The CHiME-7 Challenge: System Description and Performance of NeMo Team's DASR System

Tae Jin Park, He Huang, Ante Jukic +7

We present the NVIDIA NeMo team's multi-channel speech recognition system for the 7th CHiME Challenge Distant Automatic Speech Recognition (DASR) Task, focusing on the development…

eess.AS2023

Property-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generation

Tae Jin Park, He Huang, Coleman Hooper +5

We introduce a sophisticated multi-speaker speech data simulator, specifically engineered to generate multi-speaker speech recordings. A notable feature of this simulator is its ca…

eess.AS2023

Investigating End-to-End ASR Architectures for Long Form Audio Transcription

Nithin Rao Koluguri, Samuel Kriman, Georgy Zelenfroind +5

This paper presents an overview and evaluation of some of the end-to-end ASR models on long-form audios. We study three categories of Automatic Speech Recognition(ASR) models based…