activity
20242026
collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2026

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions

Leying Zhang, Bowen Shi, Haibin Wu +2

The rapid advancement of generative audio models has outpaced the development of robust evaluation methodologies. Existing objective metrics and general multimodal large language m…

eess.AS2025

SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation

Haitian Lu, Gaofeng Cheng, Liuping Luo +3

Recently, ``textless" speech language models (SLMs) based on speech units have made huge progress in generating naturalistic speech, including non-verbal vocalizations. However, th…

eess.AS2024

Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling

Leying Zhang, Wangyou Zhang, Chenda Li +1

Recent speech enhancement models have shown impressive performance gains by scaling up model complexity and training data. However, the impact of dataset variability (e.g. text, la…

eess.AS2024

CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations

Leying Zhang, Yao Qian, Long Zhou +9

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along w…

eess.AS2023

DDTSE: Discriminative Diffusion Model for Target Speech Extraction

Leying Zhang, Yao Qian, Linfeng Yu +5

Diffusion models have gained attention in speech enhancement tasks, providing an alternative to conventional discriminative methods. However, research on target speech extraction u…