activity
20222024
most citedAudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

21 citations · 24 across the 10 of their papers we have counts for

collaborators

10 papers

cs.SD2024

SingOMD: Singing Oriented Multi-resolution Discrete Representation Construction from Speech Models

Yuxun Tang, Yuning Wu, Jiatong Shi +1

Discrete representation has shown advantages in speech generation tasks, wherein discrete tokens are derived by discretizing hidden features from self-supervised learning (SSL) pre…

cs.SD2024

MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model

Jiatong Shi, Xutai Ma, Hirofumi Inaguma +2

Speech discrete representation has proven effective in various downstream applications due to its superior compression rate of the waveform, fast convergence during training, and c…

cs.CL2023

Joint Prediction and Denoising for Large-scale Multilingual Self-supervised Learning

William Chen, Jiatong Shi, Brian Yan +6

Multilingual self-supervised learning (SSL) has often lagged behind state-of-the-art (SOTA) methods due to the expenses and complexity required to handle many languages. This furth…

cs.CL20232 cited

Improving Cascaded Unsupervised Speech Translation with Denoising Back-translation

Yu-Kuan Fu, Liang-Hsuan Tseng, Jiatong Shi +4

Most of the speech translation models heavily rely on parallel data, which is hard to collect especially for low-resource languages. To tackle this issue, we propose to build a cas…

cs.CL202321 cited

AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Rongjie Huang, Mingze Li, Dongchao Yang +10

Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the rece…

cs.SD2023

Enhancing Speech-to-Speech Translation with Multiple TTS Targets

Jiatong Shi, Yun Tang, Ann Lee +4

It has been known that direct speech-to-speech translation (S2ST) models usually suffer from the data scarcity issue because of the limited existing parallel materials for both sou…