5 papers
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
Senqiao Yang, Kaichen Zhang, Zhaoyang Jia +20
Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process…
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Xinjie Zhang, Peng Zhang, Shicheng Zheng +21
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to…
Geometry-Aware Implicit Memory for Video World Models
Zhengxuan Wei, Xu Guo, Xinghui Li +8
Video world models aim to simulate controllable visual environments, but long-horizon rollouts depend on what the model remembers after observations leave its native context window…
ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation
Zongyi Li, Shujie Hu, Shujie Liu +7
Text-to-video models have recently undergone rapid and substantial advancements. Nevertheless, due to limitations in data and computational resources, achieving efficient generatio…
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
Zhongwei Ren, Yunchao Wei, Xun Guo +4
This work explores whether a deep generative model can learn complex knowledge solely from visual input, in contrast to the prevalent focus on text-based models like large language…