activity
20242026
most citedFoundation Models for Music: A Survey

8 citations · 8 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS2026

ARCHI-TTS: A flow-matching-based Text-to-Speech Model with Self-supervised Semantic Aligner and Accelerated Inference

Chunyat Wu, Jiajun Deng, Zhengxi Liu +3

Although diffusion-based, non-autoregressive text-to-speech (TTS) systems have demonstrated impressive zero-shot synthesis capabilities, their efficacy is still hindered by two key…

cs.SD2025

PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation

Yujia Xiao, Liumeng Xue, Lei He +8

Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into…

cs.SD2025

Inference-time Scaling for Diffusion-based Audio Super-resolution

Yizhu Jin, Zhen Ye, Zeyue Tian +4

Diffusion models have demonstrated remarkable success in generative tasks, including audio super-resolution (SR). In many applications like movie post-production and album masterin…

eess.AS2025

Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Zhen Ye, Xinfa Zhu, Chi-Min Chan +17

Recent advances in text-based large language models (LLMs), particularly in the GPT series and the o1 model, have demonstrated the effectiveness of scaling both training-time and i…

cs.SD20248 cited

Foundation Models for Music: A Survey

Yinghao Ma, Anders Øland, Anton Ragni +39

In recent years, foundation models (FMs) such as large language models (LLMs) and latent diffusion models (LDMs) have profoundly impacted diverse sectors, including music. This com…