133 citations · 170 across the 9 of their papers we have counts for
12 papers
VibeVoice-ASR-Streaming Technical Report
Yujie Tu, Zhiliang Peng, Jianwei Yu +11
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks w…
VibeVoice-ASR-BitNet Technical Report
Songchen Xu, Ting Song, Shaohan Huang +10
We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computati…
VIBEVOICE-ASR Technical Report
Zhiliang Peng, Jianwei Yu, Yaoyao Chang +21
This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation an…
VibeVoice Technical Report
Zhiliang Peng, Jianwei Yu, Wenhui Wang +10
This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeli…
Multimodal Latent Language Modeling with Next-Token Diffusion
Yutao Sun, Hangbo Bao, Wenhui Wang +5
Multimodal generative models require a unified approach to handle both discrete data (e.g., text and code) and continuous data (e.g., image, audio, video). In this work, we propose…
Kosmos-G: Generating Images in Context with Multimodal Large Language Models
Xichen Pan, Li Dong, Shaohan Huang +3
Recent advancements in subject-driven image generation have made significant strides. However, current methods still fall short in diverse application scenarios, as they require te…