4 papers
JoyStreamer: Unlocking Highly Expressive Avatars via Harmonized Text-Audio Conditioning
Ruikui Wang, Jinheng Feng, Lang Tian +6
Existing video avatar models have demonstrated impressive capabilities in scenarios such as talking, public speaking, and singing. However, the majority of these methods exhibit li…
JoyStreamer-Flash: Real-time and Infinite Audio-Driven Avatar Generation with Autoregressive Diffusion
Chaochao Li, Ruikui Wang, Liangbo Zhou +5
Existing DiT-based audio-driven avatar generation methods have achieved considerable progress, yet their broader application is constrained by limitations such as high computationa…
PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition
Li Fu, Yu Xin, Sunlu Zeng +3
This paper presents a Pronunciation-Aware Contextualized (PAC) framework to address two key challenges in Large Language Model (LLM)-based Automatic Speech Recognition (ASR) system…
UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition
Li Fu, Shanyong Yu, Siqi Li +3
Recent advancements in scaling up models have significantly improved performance in Automatic Speech Recognition (ASR) tasks. However, training large ASR models from scratch remain…