16 citations · 40 across the 16 of their papers we have counts for
5 papers · 1 filter
Fine-grained Audio-Visual Joint Representations for Multimodal Large Language Models
Guangzhi Sun, Wenyi Yu, Changli Tang +6
Audio-visual large language models (LLM) have drawn significant attention, yet the fine-grained combination of both input streams is rather under-explored, which is challenging but…
Connecting Speech Encoder and Large Language Model for ASR
Wenyi Yu, Changli Tang, Guangzhi Sun +6
The impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies a…
Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias
Ziyue Jiang, Yi Ren, Zhenhui Ye +9
Scaling text-to-speech to a large and wild dataset has been proven to be highly effective in achieving timbre and speech style generalization, particularly in zero-shot TTS. Howeve…
Leveraging phone-level linguistic-acoustic similarity for utterance-level pronunciation scoring
Wei Liu, Kaiqi Fu, Xiaohai Tian +4
Recent studies on pronunciation scoring have explored the effect of introducing phone embeddings as reference pronunciation, but mostly in an implicit manner, i.e., addition or con…
An ASR-free Fluency Scoring Approach with Self-Supervised Learning
Wei Liu, Kaiqi Fu, Xiaohai Tian +4
A typical fluency scoring system generally relies on an automatic speech recognition (ASR) system to obtain time stamps in input speech for either the subsequent calculation of flu…