8 papers · 1 filter
Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
Runjia Qian, Zile Wang, Jihai Zhang +14
Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, e…
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Hongbo Liu, Peixian Chen, Sihan Liu +12
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored.…
VGI-Bench: Probing Visual Intelligence in Video Generation Models
Xuan He, Cong Wei, Yuhao Cheng +20
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: b…
DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
Jiaxing Li, Kai Zou, Cindy Zhou +7
Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initia…
EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding
Kai Zou, Hongbo Liu, Dian Zheng +3
In this work, we present EchoGen, a unified framework for layout-to-image generation and image grounding, capable of generating images with accurate layouts and high fidelity to te…
AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
Dian Zheng, Manyuan Zhang, Hongyu Li +7
Unified multimodal models for image generation and understanding represent a significant step toward AGI and have attracted widespread attention from researchers. The main challeng…