15 papers
PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation
Qirui Li, Jinkun Hao, Yibo Li +3
Recent advances in physics-grounded video generation leverage physics simulation as a physical prior to guide video synthesis toward physically plausible outcomes. The simulation p…
Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency
Zihan Su, Teng Hu, Jiangning Zhang +4
The paper introduces Cycle-World, a framework that uses reverse‑prediction cycle consistency to reduce error accumulation in long‑horizon video generation, improving temporal consi…
UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy
Yicheng Xu, Jiangning Zhang, Zhucun Xue +5
In-context learning (ICL) enables fast task adaptation from demonstrations without per-task parameter updates but remains highly sensitive to example selection and formatting. In u…
MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data
Teng Hu, Mingchun Lu, Yating Wang +6
Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a sin…
IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction
Ran Yi, Teng Hu, Zihan Su +2
Autoregressive models have emerged as a powerful paradigm for visual content creation, but often overlook the intrinsic structural properties of visual data. Our prior work, IAR, i…
Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory
Jinzhuo Liu, Jiangning Zhang, Wencan Jiang +5
Autoregressive video generation has improved rapidly in visual fidelity and interactivity, but it still suffers from long-term inconsistency and memory degradation. Most existing s…