most citedDiscrete Audio Representation as an Alternative to Mel-Spectrograms for Speaker and Speech Recognition

2 citations · 2 across the 6 of their papers we have counts for

collaborators

7 papers

eess.AS2023

The CHiME-7 Challenge: System Description and Performance of NeMo Team's DASR System

Tae Jin Park, He Huang, Ante Jukic +7

We present the NVIDIA NeMo team's multi-channel speech recognition system for the 7th CHiME Challenge Distant Automatic Speech Recognition (DASR) Task, focusing on the development…

eess.AS2023

Property-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generation

Tae Jin Park, He Huang, Coleman Hooper +5

We introduce a sophisticated multi-speaker speech data simulator, specifically engineered to generate multi-speaker speech recordings. A notable feature of this simulator is its ca…

cs.CL2023

SALM: Speech-augmented Language Model with In-context Learning for Speech Recognition and Translation

Zhehuai Chen, He Huang, Andrei Andrusenko +6

We present a novel Speech Augmented Language Model (SALM) with {\em multitask} and {\em in-context} learning capabilities. SALM comprises a frozen text LLM, a audio encoder, a moda…

eess.AS2023

Investigating End-to-End ASR Architectures for Long Form Audio Transcription

Nithin Rao Koluguri, Samuel Kriman, Georgy Zelenfroind +5

This paper presents an overview and evaluation of some of the end-to-end ASR models on long-form audios. We study three categories of Automatic Speech Recognition(ASR) models based…

eess.AS20232 cited

Discrete Audio Representation as an Alternative to Mel-Spectrograms for Speaker and Speech Recognition

Krishna C. Puvvada, Nithin Rao Koluguri, Kunal Dhawan +2

Discrete audio representation, aka audio tokenization, has seen renewed interest driven by its potential to facilitate the application of text language modeling approaches in audio…

eess.AS2023

Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach

Tae Jin Park, Kunal Dhawan, Nithin Koluguri +1

Large language models (LLMs) have shown great promise for capturing contextual information in natural language processing tasks. We propose a novel approach to speaker diarization…