activity
20182022
most citedA spelling correction model for end-to-end speech recognition

2 citations · 6 across the 7 of their papers we have counts for

collaborators

8 papers

cs.CL2022

Biased Self-supervised learning for ASR

Florian L. Kreyssig, Yangyang Shi, Jinxi Guo +3

Self-supervised learning via masked prediction pre-training (MPPT) has shown impressive performance on a range of speech-processing tasks. This paper proposes a method to bias self…

eess.AS2022

VADOI:Voice-Activity-Detection Overlapping Inference For End-to-end Long-form Speech Recognition

Jinhan Wang, Xiaosu Tong, Jinxi Guo +2

While end-to-end models have shown great success on the Automatic Speech Recognition task, performance degrades severely when target sentences are long-form. The previous proposed…

eess.AS20202 cited

REDAT: Accent-Invariant Representation for End-to-End ASR by Domain Adversarial Training with Relabeling

Hu Hu, Xuesong Yang, Zeynab Raeesy +6

Accents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DA…

eess.AS20201 cited

Variable frame rate-based data augmentation to handle speaking-style variability for automatic speaker verification

Amber Afshan, Jinxi Guo, Soo Jin Park +3

The effects of speaking-style variability on automatic speaker verification were investigated using the UCLA Speaker Variability database which comprises multiple speaking styles p…

eess.AS20201 cited

Efficient minimum word error rate training of RNN-Transducer for end-to-end speech recognition

Jinxi Guo, Gautam Tiwari, Jasha Droppo +4

In this work, we propose a novel and efficient minimum word error rate (MWER) training method for RNN-Transducer (RNN-T). Unlike previous work on this topic, which performs on-the-…

eess.AS2019

Singing voice conversion with non-parallel data

Xin Chen, Wei Chu, Jinxi Guo +1

Singing voice conversion is a task to convert a song sang by a source singer to the voice of a target singer. In this paper, we propose using a parallel data free, many-to-one voic…