most citedEfficient Sequence Transduction by Jointly Predicting Tokens and Durations

6 citations · 10 across the 7 of their papers we have counts for

collaborators

7 papers

eess.AS2023

The CHiME-7 Challenge: System Description and Performance of NeMo Team's DASR System

Tae Jin Park, He Huang, Ante Jukic +7

We present the NVIDIA NeMo team's multi-channel speech recognition system for the 7th CHiME Challenge Distant Automatic Speech Recognition (DASR) Task, focusing on the development…

eess.AS2023

Property-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generation

Tae Jin Park, He Huang, Coleman Hooper +5

We introduce a sophisticated multi-speaker speech data simulator, specifically engineered to generate multi-speaker speech recordings. A notable feature of this simulator is its ca…

cs.CL2023

SALM: Speech-augmented Language Model with In-context Learning for Speech Recognition and Translation

Zhehuai Chen, He Huang, Andrei Andrusenko +6

We present a novel Speech Augmented Language Model (SALM) with {\em multitask} and {\em in-context} learning capabilities. SALM comprises a frozen text LLM, a audio encoder, a moda…

cs.CL2023

Leveraging Pretrained ASR Encoders for Effective and Efficient End-to-End Speech Intent Classification and Slot Filling

He Huang, Jagadeesh Balam, Boris Ginsburg

We study speech intent classification and slot filling (SICSF) by proposing to use an encoder pretrained on speech recognition (ASR) to initialize an end-to-end (E2E) Conformer-Tra…

eess.AS20236 cited

Efficient Sequence Transduction by Jointly Predicting Tokens and Durations

Hainan Xu, Fei Jia, Somshubra Majumdar +3

This paper introduces a novel Token-and-Duration Transducer (TDT) architecture for sequence-to-sequence tasks. TDT extends conventional RNN-Transducer architectures by jointly pred…

cs.CV20232 cited

Meta Compositional Referring Expression Segmentation

Li Xu, Mark He Huang, Xindi Shang +3

Referring expression segmentation aims to segment an object described by a language expression from an image. Despite the recent progress on this task, existing models tackling thi…