activity
20242026
collaborators

6 papers

cs.SD2026

DeepASMR: LLM-Based Zero-Shot ASMR Speech Generation for Anyone of Any Voice

Leying Zhang, Tingxiao Zhou, Haiyang Sun +2

While modern Text-to-Speech (TTS) systems achieve high fidelity for read-style speech, they struggle to generate Autonomous Sensory Meridian Response (ASMR), a specialized, low-int…

cs.SD2025

Training Text-to-Speech Model with Purely Synthetic Data: Feasibility, Sensitivity, and Generalization Capability

Tingxiao Zhou, Leying Zhang, Zhengyang Chen +1

The potential of synthetic data in text-to-speech (TTS) model training has gained increasing attention, yet its rationality and effectiveness require systematic validation. In this…

cs.SD2025

CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching

Leying Zhang, Yao Qian, Xiaofei Wang +8

Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing syste…

cs.SD2025

Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction

Leying Zhang, Wangyou Zhang, Zhengyang Chen +1

The acoustic background plays a crucial role in natural conversation. It provides context and helps listeners understand the environment, but a strong background makes it difficult…

eess.AS2025

SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation

Haitian Lu, Gaofeng Cheng, Liuping Luo +3

Recently, ``textless" speech language models (SLMs) based on speech units have made huge progress in generating naturalistic speech, including non-verbal vocalizations. However, th…

eess.AS2024

Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling

Leying Zhang, Wangyou Zhang, Chenda Li +1

Recent speech enhancement models have shown impressive performance gains by scaling up model complexity and training data. However, the impact of dataset variability (e.g. text, la…