papers

Publications (27)

cs.CL2025

Fewer Hallucinations, More Verification: A Three-Stage LLM-Based Framework for ASR Error Correction

Yangui Fang, Baixu Chen, Jing Peng +4

Automatic Speech Recognition (ASR) error correction aims to correct recognition errors while preserving accurate text. Although traditional approaches demonstrate moderate effectiv…

eess.AS2024

NTC-KWS: Noise-aware CTC for Robust Keyword Spotting

Yu Xi, Haoyu Li, Hao Li +4

In recent years, there has been a growing interest in designing small-footprint yet effective Connectionist Temporal Classification based keyword spotting (CTC-KWS) systems. They a…

eess.AS2024

Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario

Wen Wen, Qiang Zhou, Yu Xi +3

In multi-speaker scenarios, leveraging spatial features is essential for enhancing target speech. While with limited microphone arrays, developing a compact multi-channel speech en…

eess.AS2024

Text-aware Speech Separation for Multi-talker Keyword Spotting

Haoyu Li, Baochen Yang, Yu Xi +4

For noisy environments, ensuring the robustness of keyword spotting (KWS) systems is essential. While much research has focused on noisy KWS, less attention has been paid to multi-…

eess.AS2026

TASU: Text-Only Alignment for Speech Understanding

Jing Peng, Yi Yang, Xu Li +5

Recent advances in Speech Large Language Models (Speech LLMs) have paved the way for unified architectures across diverse speech understanding tasks. However, prevailing alignment…

eess.AS2025

Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning

Yangui Fang, Jing Peng, Xu Li +4

Recent advances in automatic speech recognition (ASR) have combined speech encoders with large language models (LLMs) through projection, forming Speech LLMs with strong performanc…