5 papers
Diagnose, Then Refine: A Closed-Loop TTS System with AudioLLM-Guided Correction
Zeyang Song, Tianchi Liu, Tianrui Wang +3
Current TTS systems typically rely on open-loop, single-pass generation and can produce sporadic local prosodic defects, such as misplaced stress, unnatural pauses, or flattened in…
EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis
Tianchi Liu, Zeyang Song, Tianrui Wang +3
Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotion…
Long-Context Modeling Networks for Monaural Speech Enhancement: A Comparative Study
Qiquan Zhang, Moran Chen, Zeyang Song +3
Advanced long-context modeling backbone networks, such as Transformer, Conformer, and Mamba, have demonstrated state-of-the-art performance in speech enhancement. However, a system…
IML-Spikeformer: Input-aware Multi-Level Spiking Transformer for Speech Processing
Zeyang Song, Shimin Zhang, Yuhong Chou +2
Spiking Neural Networks (SNNs), inspired by biological neural mechanisms, represent a promising neuromorphic computing paradigm that offers energy-efficient alternatives to traditi…
ED-sKWS: Early-Decision Spiking Neural Networks for Rapid,and Energy-Efficient Keyword Spotting
Zeyang Song, Qianhui Liu, Qu Yang +2
Keyword Spotting (KWS) is essential in edge computing requiring rapid and energy-efficient responses. Spiking Neural Networks (SNNs) are well-suited for KWS for their efficiency an…