most citedMega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

16 citations · 40 across the 16 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS20235 cited

Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models

Guangzhi Sun, Wenyi Yu, Changli Tang +6

Audio-visual large language models (LLM) have drawn significant attention, yet the fine-grained combination of both input streams is rather under-explored, which is challenging but…

eess.AS2023

Connecting Speech Encoder and Large Language Model for ASR

Wenyi Yu, Changli Tang, Guangzhi Sun +6

The impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies a…

eess.AS202316 cited

Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Ziyue Jiang, Yi Ren, Zhenhui Ye +9

Scaling text-to-speech to a large and wild dataset has been proven to be highly effective in achieving timbre and speech style generalization, particularly in zero-shot TTS. Howeve…

eess.AS2023

Leveraging phone-level linguistic-acoustic similarity for utterance-level pronunciation scoring

Wei Liu, Kaiqi Fu, Xiaohai Tian +4

Recent studies on pronunciation scoring have explored the effect of introducing phone embeddings as reference pronunciation, but mostly in an implicit manner, i.e., addition or con…

eess.AS2023

An ASR-free Fluency Scoring Approach with Self-Supervised Learning

Wei Liu, Kaiqi Fu, Xiaohai Tian +4

A typical fluency scoring system generally relies on an automatic speech recognition (ASR) system to obtain time stamps in input speech for either the subsequent calculation of flu…