activity
20222024
most citedNaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

20 citations · 49 across the 13 of their papers we have counts for

collaborators

13 papers

cs.CL20242 cited

TCMD: A Traditional Chinese Medicine QA Dataset for Evaluating Large Language Models

Ping Yu, Kaitao Song, Fengchen He +2

The recently unprecedented advancements in Large Language Models (LLMs) have propelled the medical community by establishing advanced medical-domain models. However, due to the lim…

eess.AS202420 cited

NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Zeqian Ju, Yuancheng Wang, Kai Shen +16

While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intric…

eess.SP20247 cited

EEGFormer: Towards Transferable and Interpretable Large-Scale EEG Foundation Model

Yuqi Chen, Kan Ren, Kaitao Song +4

Self-supervised learning has emerged as a highly effective approach in the fields of natural language processing and computer vision. It is also applicable to brain signals such as…

cs.CL20244 cited

EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction

Siyu Yuan, Kaitao Song, Jiangjie Chen +5

To address intricate real-world tasks, there has been a rising interest in tool utilization in applications of large language models (LLMs). To develop LLM-based agents, it usually…

cs.CL20232 cited

MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models

Dingyao Yu, Kaitao Song, Peiling Lu +5

AI-empowered music processing is a diverse field that encompasses dozens of tasks, ranging from generation tasks (e.g., timbre synthesis) to comprehension tasks (e.g., music classi…

eess.AS2023

PromptTTS 2: Describing and Generating Voices with Text Prompt

Yichong Leng, Zhifang Guo, Kai Shen +12

Speech conveys more information than text, as the same word can be uttered in various voices to convey diverse information. Compared to traditional text-to-speech (TTS) methods rel…