activity
20242026
collaborators

8 papers

cs.CV2026

SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models

Zizhao Tong, Yeying Jin, Hongfeng Lai +11

Interactive world models for first-person shooter (FPS) games must resolve high-frequency overlapping control signals at every frame without disrupting unaffected regions. Existing…

cs.CV2026

TIE: Time Interval Encoding for Video Generation over Events

Zhilei Shu, Shangwen Zhu, Zihang Liang +10

Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which 68% of general clips and ove…

cs.CV2026

Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models

Shangwen Zhu, Qianyu Peng, Zhao Pu +12

Modern interactive video world models have achieved impressive visual fidelity, yet lack fine-grained multi-entity control and cross-entity, cross-world generalization. We trace th…

cs.CV2025

Addressing the ID-Matching Challenge in Long Video Captioning

Zhantao Yang, Huangji Wang, Ruili Feng +6

Generating captions for long and complex videos is both critical and challenging, with significant implications for the growing fields of text-to-video generation and multi-modal u…

cs.CV2025

MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance

Shangwen Zhu, Qianyu Peng, Zhilei Shu +9

High-fidelity text-to-image and text-to-video generation typically relies on Classifier-Free Guidance (CFG), but achieving optimal results often demands computationally expensive s…

cs.LG2025

Instability in Diffusion ODEs: An Explanation for Inaccurate Image Reconstruction

Han Zhang, Jinghong Mao, Shangwen Zhu +6

Diffusion reconstruction plays a critical role in various applications such as image editing, restoration, and style transfer. In theory, the reconstruction should be simple - it j…