21 citations · 24 across the 10 of their papers we have counts for
10 papers
SingOMD: Singing Oriented Multi-resolution Discrete Representation Construction from Speech Models
Yuxun Tang, Yuning Wu, Jiatong Shi +1
Discrete representation has shown advantages in speech generation tasks, wherein discrete tokens are derived by discretizing hidden features from self-supervised learning (SSL) pre…
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
Jiatong Shi, Xutai Ma, Hirofumi Inaguma +2
Speech discrete representation has proven effective in various downstream applications due to its superior compression rate of the waveform, fast convergence during training, and c…
Joint Prediction and Denoising for Large-scale Multilingual Self-supervised Learning
William Chen, Jiatong Shi, Brian Yan +6
Multilingual self-supervised learning (SSL) has often lagged behind state-of-the-art (SOTA) methods due to the expenses and complexity required to handle many languages. This furth…
Improving Cascaded Unsupervised Speech Translation with Denoising Back-translation
Yu-Kuan Fu, Liang-Hsuan Tseng, Jiatong Shi +4
Most of the speech translation models heavily rely on parallel data, which is hard to collect especially for low-resource languages. To tackle this issue, we propose to build a cas…
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
Rongjie Huang, Mingze Li, Dongchao Yang +10
Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the rece…
Enhancing Speech-to-Speech Translation with Multiple TTS Targets
Jiatong Shi, Yun Tang, Ann Lee +4
It has been known that direct speech-to-speech translation (S2ST) models usually suffer from the data scarcity issue because of the limited existing parallel materials for both sou…