activity
20242026
collaborators

14 papers

cs.CV2026

MOSS-VL Technical Report

Pengyu Wang, Chenkun Tan, Shaojun Zhou +29

We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across th…

eess.AS2026

One-Step Token-to-Waveform Generation with MeanFlow in Latent Space

Zheqi Dai, Guangyan Zhang, Zhen Ye +5

Neural audio codecs are central to modern LLM-based Text-to-Speech (TTS) and multimodal systems. As low-bitrate semantic codecs gain prominence, the Token-to-Waveform (Token2Wav) d…

eess.AS2026

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation

Zhen Ye, Xu Tan, Yiming Li +10

Spoken dialogue models typically start from text LLM backbones, yet reasoning often degrades when conditioning on speech instead of text. We attribute part of this modality gap to…

eess.AS2026

MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information Disentanglement

Jingyu Li, Guangyan Zhang, Zhen Ye +1

Audio codecs are a critical component of modern speech generation systems. This paper introduces a low-bitrate, multi-scale residual codec that encodes speech into four distinct st…

cs.CL2025

RL from Teacher-Model Refinement: Gradual Imitation Learning for Machine Translation

Dongyub Jude Lee, Zhenyi Ye, Pengcheng He

Preference-learning methods for machine translation (MT), such as Direct Preference Optimization (DPO), have shown strong gains but typically rely on large, carefully curated prefe…

eess.AS2025

YuE: Scaling Open Foundation Models for Long-Form Music Generation

Ruibin Yuan, Hanfeng Lin, Shuyue Guo +55

We tackle the task of long-form music generation--particularly the challenging \textbf{lyrics-to-song} problem--by introducing YuE, a family of open foundation models based on the…