4 papers
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
Yupeng Zhou, Lianghua Huang, Zhifan Wu +7
In this work, we propose Mutual Forcing, a framework for fast autoregressive audio-video generation with long-horizon audio-video synchronization. Our approach addresses two key ch…
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
Yupeng Zhou, Zhen Li, Ziheng Ouyang +8
Encoding videos into discrete tokens could align with text tokens to facilitate concise and unified multi-modal LLMs, yet introducing significant spatiotemporal compression compare…
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
Xuying Zhang, Yupeng Zhou, Kai Wang +6
Novel view synthesis (NVS) is a cornerstone for image-to-3d creation. However, existing works still struggle to maintain consistency between the generated views and the input views…
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
Yi Xin, Juncheng Yan, Qi Qin +18
We present Lumina-mGPT 2.0, a stand-alone, decoder-only autoregressive model that revisits and revitalizes the autoregressive paradigm for high-quality image generation and beyond.…