◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Dongchao Yang

24 papers hereh-index 131.3k citations32 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author1
  • first author6
  • middle author16

Across the 23 of 24 papers where every author was matched, so the position is known.

fields
  • cs.SD14
  • eess.AS9
  • cs.CL1
same name
  • Dongchao Yang — 16 papers, h 16
  • Dongchao Yang — 8 papers, h 3
  • Dongchao Yang — 7 papers
  • Dongchao Yang — 1 paper

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20232026
most citedNaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

20 citations · 25 across the 20 of their papers we have counts for

collaborators
Showing eess.ASShow all

4 papers · 1 filter

eess.AS2025★ 2 cited

Kimi-Audio Technical Report

KimiTeam, Ding Ding, Zeqian Ju +37

We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, inclu…

eess.AS2025

MoonCast: High-Quality Zero-Shot Podcast Generation

Zeqian Ju, Dongchao Yang, Jianwei Yu +7

Recent advances in text-to-speech synthesis have achieved notable success in generating high-quality short utterances for individual speakers. However, these systems still face cha…

eess.AS2024

Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models

Haibin Wu, Xuanjun Chen, Yi-Cheng Lin +13

Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The i…

eess.AS2024

AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions

Yuanyuan Wang, Hangting Chen, Dongchao Yang +2

Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.