6 papers
VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation
Shikun Sun, Liao Qu, Huichao Zhang +8
Visual generation is dominated by three paradigms: AutoRegressive (AR), diffusion, and Visual AutoRegressive (VAR) models. Unlike AR and diffusion, VARs operate on heterogeneous in…
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
Huichao Zhang, Liao Qu, Yiheng Liu +33
We present NextFlow, a unified decoder-only autoregressive transformer trained on 6 trillion interleaved text-image discrete tokens. By leveraging a unified vision representation w…
Unveiling the Attribute Misbinding Threat in Identity-Preserving Models
Junming Fu, Jishen Zeng, Yi Jiang +4
Identity-preserving models have led to notable progress in generating personalized content. Unfortunately, such models also exacerbate risks when misused, for instance, by generati…
Advancing the Foundation Model for Music Understanding
Yi Jiang, Wei Wang, Xianwen Guo +6
The field of Music Information Retrieval (MIR) is fragmented, with specialized models excelling at isolated tasks. In this work, we challenge this paradigm by introducing a unified…
Better Recommendations: Validating AI-generated Subject Terms Through LOC Linked Data Service
Kwok Leong Tang, Yi Jiang
This article explores the integration of AI-generated subject terms into library cataloging, focusing on validation through the Library of Congress Linked Data Service. It examines…
AI TrackMate: Finally, Someone Who Will Give Your Music More Than Just "Sounds Great!"
Yi-Lin Jiang, Chia-Ho Hsiung, Yen-Tung Yeh +2
The rise of "bedroom producers" has democratized music creation, while challenging producers to objectively evaluate their work. To address this, we present AI TrackMate, an LLM-ba…