activity
20242026
collaborators

5 papers

cs.SD2026

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

Hankun Wang, Bohan Li, Shi Lian +6

Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface fo…

eess.AS2026

A Unified and Reproducible Experimentation Framework for Speech Understanding

Jing Peng, Junhao Du, Chenghao Wang +21

Speech foundation models and Speech LLMs have advanced speech understanding, yet deployment-oriented model selection is hindered by non-comparable evaluations caused by mismatched…

eess.AS2025

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding

Yu Xi, Haoyu Li, Xiaoyu Gu +2

Keyword spotting (KWS) is essential for voice-driven applications, demanding both accuracy and efficiency. Traditional ASR-based KWS methods, such as greedy and beam search, explor…

cs.SD2025

Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding

Yu Xi, Xiaoyu Gu, Haoyu Li +3

RNN-T-based keyword spotting (KWS) with autoregressive decoding~(AR) has gained attention due to its streaming architecture and superior performance. However, the simplicity of the…

eess.AS2024

Streaming Keyword Spotting Boosted by Cross-layer Discrimination Consistency

Yu Xi, Haoyu Li, Xiaoyu Gu +3

Connectionist Temporal Classification (CTC), a non-autoregressive training criterion, is widely used in online keyword spotting (KWS). However, existing CTC-based KWS decoding stra…