papers

Publications (77)

cs.CL2024

Comparing Discrete and Continuous Space LLMs for Speech Recognition

Yaoxun Xu, Shi-Xiong Zhang, Jianwei Yu +2

This paper investigates discrete and continuous speech representations in Large Language Model (LLM)-based Automatic Speech Recognition (ASR), organizing them by feature continuity…

eess.AS2025

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

Yifan Yang, Shujie Liu, Jinyu Li +10

This paper introduces Interleaved Speech-Text Language Model (IST-LM) for zero-shot streaming Text-to-Speech (TTS). Unlike many previous approaches, IST-LM is directly trained on i…

cs.SD2022

Investigation of Data Augmentation Techniques for Disordered Speech Recognition

Mengzhe Geng, Xurong Xie, Shansong Liu +4

Disordered speech recognition is a highly challenging task. The underlying neuro-motor conditions of people with speech disorders, often compounded with co-occurring physical disab…

cs.SD2024

Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings

He Zhao, Hangting Chen, Jianwei Yu +1

Target speaker extraction (TSE) aims to extract the target speaker's voice from the input mixture. Previous studies have concentrated on high-overlapping scenarios. However, real-w…

eess.AS2023

High Fidelity Speech Enhancement with Band-split RNN

Jianwei Yu, Yi Luo, Hangting Chen +2

Despite the rapid progress in speech enhancement (SE) research, enhancing the quality of desired speech in environments with strong noise and interfering speakers remains challengi…

eess.AS2022

Neural Architecture Search For LF-MMI Trained Time Delay Neural Networks

Shoukang Hu, Xurong Xie, Mingyu Cui +6

State-of-the-art automatic speech recognition (ASR) system development is data and computation intensive. The optimal design of deep neural networks (DNNs) for these systems often…