activity
20202026
most citedAdapting GPT, GPT-2 and BERT Language Models for Speech Recognition

5 citations · 9 across the 11 of their papers we have counts for

collaborators

11 papers

cs.CL2026

Streaming Speech-to-Text Translation with a SpeechLLM

Titouan Parcollet, Shucong Zhang, Xianrui Zheng +1

Normally, a system that translates speech into text consists of separate modules for speech recognition and text-to-text translation. Combining those tasks into a SpeechLLM promise…

eess.AS2026

Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations

Xin Guo, Chunrui Zhao, Hong Jia +4

Integrating Federated Learning (FL) with self-supervised learning (SSL) enables privacy-preserving fine-tuning for speech tasks. However, federated environments exhibit significant…

eess.AS2025

DNCASR: End-to-End Training for Speaker-Attributed ASR

Xianrui Zheng, Chao Zhang, Philip C. Woodland

This paper introduces DNCASR, a novel end-to-end trainable system designed for joint neural speaker clustering and automatic speech recognition (ASR), enabling speaker-attributed t…

eess.AS2024

SOT Triggered Neural Clustering for Speaker Attributed ASR

Xianrui Zheng, Guangzhi Sun, Chao Zhang +1

This paper introduces a novel approach to speaker-attributed ASR transcription using a neural clustering method. With a parallel processing mechanism, diarisation and ASR can be ap…

eess.AS2023★ 2 cited

Conditional Diffusion Model for Target Speaker Extraction

Theodor Nguyen, Guangzhi Sun, Xianrui Zheng +2

We propose DiffSpEx, a generative target speaker extraction method based on score-based generative modelling through stochastic differential equations. DiffSpEx deploys a continuou…

cs.CL2023

Can Contextual Biasing Remain Effective with Whisper and GPT-2?

Guangzhi Sun, Xianrui Zheng, Chao Zhang +1

End-to-end automatic speech recognition (ASR) and large language models, such as Whisper and GPT-2, have recently been scaled to use vast amounts of training data. Despite the larg…