activity
20242026
collaborators

7 papers

cs.SD2026

Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer

Shengfan Shen, Di Wu, Xingchen Song +5

Expressive speech synthesis for voice assistants requires flexible style control that adapts to explicit requests and broader interaction context. We propose Harness TTS, a lightwe…

cs.SD2026

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

Dinghao Zhou, Xingchen Song, Di Wu +3

Continuous audio autoencoders reconstruct waveforms well but often produce latents with weak structure for understanding, while self-supervised audio encoders capture semantics but…

cs.SD2026

Borderless Long Speech Synthesis

Xingchen Song, Di Wu, Dinghao Zhou +12

Most existing text-to-speech (TTS) systems either synthesize speech sentence by sentence and stitch the results together, or drive synthesis from plain-text dialogues alone. Both a…

cs.SD2026

Iterate to Differentiate: Enhancing Discriminability and Reliability in Zero-Shot TTS Evaluation

Shengfan Shen, Di Wu, Xingchen Song +5

Reliable evaluation of modern zero-shot text-to-speech (TTS) models remains challenging. Subjective tests are costly and hard to reproduce, while objective metrics often saturate,…

cs.SD2025

Back to Ear: Perceptually Driven High Fidelity Music Reconstruction

Kangdi Wang, Zhiyue Wu, Dinghao Zhou +3

Variational Autoencoders (VAEs) are essential for large-scale audio tasks like diffusion-based generation. However, existing open-source models often neglect auditory perceptual as…

eess.AS2024

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

Xingchen Song, Chengdong Liang, Binbin Zhang +9

Large Automatic Speech Recognition (ASR) models demand a vast number of parameters, copious amounts of data, and significant computational resources during the training process. Ho…