7 papers
MOSS Transcribe Diarize Technical Report
MOSI. AI, :, Donghua Yu +23
Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meet…
VOGUE: A Multimodal Dataset for Conversational Recommendation in Fashion
David Guo, Minqi Sun, Yilun Jiang +2
Multimodal conversational recommendation has recently emerged as a promising paradigm for delivering personalized experiences through natural dialogue enriched by visual and contex…
VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation
Shikun Sun, Liao Qu, Huichao Zhang +8
Visual generation is dominated by three paradigms: AutoRegressive (AR), diffusion, and Visual AutoRegressive (VAR) models. Unlike AR and diffusion, VARs operate on heterogeneous in…
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
Huichao Zhang, Liao Qu, Yiheng Liu +33
We present NextFlow, a unified decoder-only autoregressive transformer trained on 6 trillion interleaved text-image discrete tokens. By leveraging a unified vision representation w…
Unveiling the Attribute Misbinding Threat in Identity-Preserving Models
Junming Fu, Jishen Zeng, Yi Jiang +4
Identity-preserving models have led to notable progress in generating personalized content. Unfortunately, such models also exacerbate risks when misused, for instance, by generati…
Advancing the Foundation Model for Music Understanding
Yi Jiang, Wei Wang, Xianwen Guo +6
The field of Music Information Retrieval (MIR) is fragmented, with specialized models excelling at isolated tasks. In this work, we challenge this paradigm by introducing a unified…