activity
20182024
most citedPrompting Large Language Models with Speech Recognition Abilities

2 citations · 9 across the 10 of their papers we have counts for

collaborators
Showing eess.ASShow all

10 papers · 1 filter

eess.AS2024

Effective internal language model training and fusion for factorized transducer model

Jinxi Guo, Niko Moritz, Yingyi Ma +6

The internal language model (ILM) of the neural transducer has been widely studied. In most prior work, it is mainly used for estimating the ILM score and is subsequently subtracte…

eess.AS2023

Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model

Jiamin Xie, Ke Li, Jinxi Guo +7

Neural network pruning offers an effective method for compressing a multilingual automatic speech recognition (ASR) model with minimal performance loss. However, it entails several…

eess.AS20232 cited

Prompting Large Language Models with Speech Recognition Abilities

Yassir Fathullah, Chunyang Wu, Egor Lakomkin +9

Large language models have proven themselves highly flexible, able to solve a wide range of generative tasks, such as abstractive summarization and open-ended question answering. I…

eess.AS2022

VADOI:Voice-Activity-Detection Overlapping Inference For End-to-end Long-form Speech Recognition

Jinhan Wang, Xiaosu Tong, Jinxi Guo +2

While end-to-end models have shown great success on the Automatic Speech Recognition task, performance degrades severely when target sentences are long-form. The previous proposed…

eess.AS20202 cited

REDAT: Accent-Invariant Representation for End-to-End ASR by Domain Adversarial Training with Relabeling

Hu Hu, Xuesong Yang, Zeynab Raeesy +6

Accents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DA…

eess.AS20201 cited

Variable frame rate-based data augmentation to handle speaking-style variability for automatic speaker verification

Amber Afshan, Jinxi Guo, Soo Jin Park +3

The effects of speaking-style variability on automatic speaker verification were investigated using the UCLA Speaker Variability database which comprises multiple speaking styles p…