activity
20212026
most citedKosmos-2: Grounding Multimodal Large Language Models to the World

133 citations · 170 across the 9 of their papers we have counts for

collaborators

12 papers

eess.AS2026

VibeVoice-ASR-Streaming Technical Report

Yujie Tu, Zhiliang Peng, Jianwei Yu +11

Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks w…

cs.SD2026

VibeVoice-ASR-BitNet Technical Report

Songchen Xu, Ting Song, Shaohan Huang +10

We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computati…

cs.SD2026

VIBEVOICE-ASR Technical Report

Zhiliang Peng, Jianwei Yu, Yaoyao Chang +21

This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation an…

cs.CL20251 cited

VibeVoice Technical Report

Zhiliang Peng, Jianwei Yu, Wenhui Wang +10

This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeli…

cs.CL20241 cited

Multimodal Latent Language Modeling with Next-Token Diffusion

Yutao Sun, Hangbo Bao, Wenhui Wang +5

Multimodal generative models require a unified approach to handle both discrete data (e.g., text and code) and continuous data (e.g., image, audio, video). In this work, we propose…

cs.CV2023

Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Xichen Pan, Li Dong, Shaohan Huang +3

Recent advancements in subject-driven image generation have made significant strides. However, current methods still fall short in diverse application scenarios, as they require te…