102 citations · 118 across the 12 of their papers we have counts for
Showing cs.SDShow all
3 papers · 1 filter
cs.SD2026
LLaDA-TTS: Unifying Speech Synthesis and Zero-Shot Editing via Masked Diffusion Modeling
Xiaoyu Fan, Huizhi Xie, Wei Zou +1
Large language model (LLM)-based text-to-speech (TTS) systems achieve remarkable naturalness via autoregressive (AR) decoding, but require N sequential steps to generate N speech t…
cs.SD2022
Audio-Visual Wake Word Spotting System For MISP Challenge 2021
Yanguang Xu, Jianwei Sun, Yang Han +7
This paper presents the details of our system designed for the Task 1 of Multimodal Information Based Speech Processing (MISP) Challenge 2021. The purpose of Task 1 is to leverage…
cs.SD2019
Cross-task pre-training for on-device acoustic scene classification
Ruixiong Zhang, Wei Zou, Xiangang Li
Acoustic scene classification (ASC) and acoustic event detection (AED) are different but related tasks. Acoustic events can provide useful information for recognizing acoustic scen…