most citedA CTC Alignment-based Non-autoregressive Transformer for End-to-end Automatic Speech Recognition

32 citations · 64 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2025

OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics

Wei Chu, Yuanzhe Dong, Ke Tan +7

OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podca…

eess.AS2024

SOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation for Low Resource ASR

Natarajan Balaji Shankar, Ruchao Fan, Abeer Alwan

Recently, speech foundation models have gained popularity due to their superiority in finetuning downstream ASR tasks. However, models finetuned on certain domains, such as LibriSp…

eess.AS2024

Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models

Ruchao Fan, Natarajan Balaji Shankar, Abeer Alwan

Speech foundation models (SFMs) have achieved state-of-the-art results for various speech tasks in supervised (e.g. Whisper) or self-supervised systems (e.g. WavLM). However, the p…

eess.AS20242 cited

UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models

Ruchao Fan, Natarajan Balaji Shanka, Abeer Alwan

Non-autoregressive automatic speech recognition (NASR) models have gained attention due to their parallelism and fast inference. The encoder-based NASR, e.g. connectionist temporal…

eess.AS202330 cited

Towards Better Domain Adaptation for Self-supervised Models: A Case Study of Child ASR

Ruchao Fan, Yunzheng Zhu, Jinhan Wang +1

Recently, self-supervised learning (SSL) from unlabelled speech data has gained increased attention in the automatic speech recognition (ASR) community. Typical SSL methods include…

cs.CL202332 cited

A CTC Alignment-based Non-autoregressive Transformer for End-to-end Automatic Speech Recognition

Ruchao Fan, Wei Chu, Peng Chang +1

Recently, end-to-end models have been widely used in automatic speech recognition (ASR) systems. Two of the most representative approaches are connectionist temporal classification…