collaborators

7 papers

eess.AS2025

Joint decoding method for controllable contextual speech recognition based on Speech LLM

Yangui Fang, Jing Peng, Yu Xi +5

Contextual speech recognition refers to the ability to identify preferences for specific content based on contextual information. Recently, leveraging the contextual understanding…

eess.AS2025

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding

Yu Xi, Haoyu Li, Xiaoyu Gu +2

Keyword spotting (KWS) is essential for voice-driven applications, demanding both accuracy and efficiency. Traditional ASR-based KWS methods, such as greedy and beam search, explor…

eess.AS2025

Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning

Yangui Fang, Jing Peng, Xu Li +4

Recent advances in automatic speech recognition (ASR) have combined speech encoders with large language models (LLMs) through projection, forming Speech LLMs with strong performanc…

cs.SD2025

Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding

Yu Xi, Xiaoyu Gu, Haoyu Li +3

RNN-T-based keyword spotting (KWS) with autoregressive decoding~(AR) has gained attention due to its streaming architecture and superior performance. However, the simplicity of the…

eess.AS2024

Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario

Wen Wen, Qiang Zhou, Yu Xi +3

In multi-speaker scenarios, leveraging spatial features is essential for enhancing target speech. While with limited microphone arrays, developing a compact multi-channel speech en…

eess.AS2024

Streaming Keyword Spotting Boosted by Cross-layer Discrimination Consistency

Yu Xi, Haoyu Li, Xiaoyu Gu +3

Connectionist Temporal Classification (CTC), a non-autoregressive training criterion, is widely used in online keyword spotting (KWS). However, existing CTC-based KWS decoding stra…