activity
20242026
collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2026

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies

Dongchao Yang

Every audio generative system makes two coupled decisions: what representation to generate, and how to model its distribution. This paper organizes audio generative modeling around…

eess.AS2025

Kimi-Audio Technical Report

KimiTeam, Ding Ding, Zeqian Ju +37

We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, inclu…

eess.AS2025

AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions

Yuanyuan Wang, Hangting Chen, Dongchao Yang +2

Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content…

eess.AS2025

MoonCast: High-Quality Zero-Shot Podcast Generation

Zeqian Ju, Dongchao Yang, Jianwei Yu +7

Recent advances in text-to-speech synthesis have achieved notable success in generating high-quality short utterances for individual speakers. However, these systems still face cha…

eess.AS2024

Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models

Haibin Wu, Xuanjun Chen, Yi-Cheng Lin +13

Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The i…

eess.AS2024

RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Detai Xin, Xu Tan, Kai Shen +8

We present RALL-E, a robust language modeling method for text-to-speech (TTS) synthesis. While previous work based on large language models (LLMs) shows impressive performance on z…