3 papers
cs.SD2026
VIBEVOICE-ASR Technical Report
Zhiliang Peng, Jianwei Yu, Yaoyao Chang +21
This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation an…
cs.CL2025
VibeVoice Technical Report
Zhiliang Peng, Jianwei Yu, Wenhui Wang +10
This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeli…
cs.CL2024
Multimodal Latent Language Modeling with Next-Token Diffusion
Yutao Sun, Hangbo Bao, Wenhui Wang +5
Multimodal generative models require a unified approach to handle both discrete data (e.g., text and code) and continuous data (e.g., image, audio, video). In this work, we propose…