13 papers
Drift-AR: Single-Step Visual Autoregressive Generation via Anti-Symmetric Drifting
Zhen Zou, Xiaoxiao Ma, Mingde Yao +3
Autoregressive (AR)-Diffusion hybrid paradigms combine AR's structured semantic modeling with diffusion's high-fidelity synthesis, yet suffer from a dual speed bottleneck: the sequ…
AnchorEdit: Maintaining Temporal Consistency in Multi-turn Image Editing via Causal Memory
Hang Xu, Xiaoxiao Ma, Guohui Zhang +7
Multi-turn image editing is essential for iterative design, yet current models often struggle with identity drift and error accumulation over successive steps. While existing resea…
SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation
Tianfei Ren, Zhipeng Yan, Yiming Zhao +13
While text-to-image models have made strong progress in visual fidelity, faithfully realizing complex visual intents remains challenging because many requirements must be tracked a…
MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation
Xiaoxiao Ma, Jiachen Lei, Tianfei Ren +6
Reinforcement learning (RL) has been successfully applied to autoregressive (AR) and diffusion models. However, extending RL to hybrid AR-diffusion frameworks remains challenging d…
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
Naishan Zheng, Jie Huang, Qingpei Guo +1
Understanding long videos with multimodal large language models (MLLMs) remains challenging due to the heavy redundancy across frames and the need for temporally coherent represent…
Fast-ARDiff: An Entropy-informed Acceleration Framework for Continuous Space Autoregressive Generation
Zhen Zou, Xiaoxiao Ma, Jie Huang +2
Autoregressive(AR)-diffusion hybrid paradigms combine AR's structured modeling with diffusion's photorealistic synthesis, yet suffer from high latency due to sequential AR generati…