activity
20242026
collaborators

9 papers

cs.SD2026

MOSS Transcribe Diarize Technical Report

MOSI. AI, :, Donghua Yu +23

Speaker-Attributed, Time-Stamped Transcription (SATS) aims to transcribe what is said and to precisely determine the timing of each speaker, which is particularly valuable for meet…

cs.IR2026

VOGUE: A Multimodal Dataset for Conversational Recommendation in Fashion

David Guo, Minqi Sun, Yilun Jiang +2

Multimodal conversational recommendation has recently emerged as a promising paradigm for delivering personalized experiences through natural dialogue enriched by visual and contex…

cs.CV2026

VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation

Shikun Sun, Liao Qu, Huichao Zhang +8

Visual generation is dominated by three paradigms: AutoRegressive (AR), diffusion, and Visual AutoRegressive (VAR) models. Unlike AR and diffusion, VARs operate on heterogeneous in…

cs.CV2026

NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation

Huichao Zhang, Liao Qu, Yiheng Liu +33

We present NextFlow, a unified decoder-only autoregressive transformer trained on 6 trillion interleaved text-image discrete tokens. By leveraging a unified vision representation w…

cs.CR2025

Unveiling the Attribute Misbinding Threat in Identity-Preserving Models

Junming Fu, Jishen Zeng, Yi Jiang +4

Identity-preserving models have led to notable progress in generating personalized content. Unfortunately, such models also exacerbate risks when misused, for instance, by generati…

cs.SD2025

Advancing the Foundation Model for Music Understanding

Yi Jiang, Wei Wang, Xianwen Guo +6

The field of Music Information Retrieval (MIR) is fragmented, with specialized models excelling at isolated tasks. In this work, we challenge this paradigm by introducing a unified…