activity
20242026
collaborators
Showing cs.SDShow all

9 papers · 1 filter

cs.SD2026

AnyBand: Unified Multi-Bandwidth Speech Extension via Frequency-Aware In-Context Spectral Infilling

Junchuan Zhao, Minh Duc Vu, Bowen Zhang +1

Bandwidth extension (BWE) aims to recover missing high-frequency content from band-limited speech. Existing methods often formulate BWE as a fixed or predefined bandwidth conversio…

cs.SD2026

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction

Xintong Wang, Ye Wang

Accent text-to-speech (TTS) aims to synthesize speech with target accents. Existing accent TTS systems typically rely on a two-stage pipeline that first converts standard phone seq…

cs.SD2026

MSpoofTTS: Multi-Resolution Spoof-Guided Inference for Discrete Speech Synthesis

Junchuan Zhao, Minh Duc Vu, Ye Wang

Neural codec language models enable high-quality discrete speech synthesis, yet their inference remains vulnerable to token-level artifacts and distributional drift that degrade pe…

cs.SD2026

TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis

Qifan Liang, Yuansen Liu, Ruixin Wei +3

While controllable Text-to-Speech (TTS) has achieved notable progress, most existing methods remain limited to inter-utterance-level control, making fine-grained intra-utterance ex…

cs.SD2026

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance

Junchuan Zhao, Wei Zeng, Tianle Lyu +1

Singing Voice Synthesis (SVS) aims to generate expressive vocal performances from structured musical inputs such as lyrics and pitch sequences. While recent progress in discrete co…

cs.SD2026

CodecFlow: Efficient Bandwidth Extension via Conditional Flow Matching in Neural Codec Latent Space

Bowen Zhang, Junchuan Zhao, Ian McLoughlin +2

Speech Bandwidth Extension improves clarity and intelligibility by restoring/inferring appropriate high-frequency content for low-bandwidth speech. Existing methods often rely on s…