20 papers · 1 filter
Visual Geometry Transformer in the Wild: Distractor-Free 3D Reconstruction
Tianbo Pan, Xingyi Yang, Shizun Wang +1
Current end-to-end multi-view 3D reconstruction methods achieve impressive results, but rely on a restrictive static assumption: the scenes is entire distractor-free with perfect c…
BadWorld: Adversarial Attacks on World Models
Linghui Shen, Mingyue Cui, Xingyi Yang
Visual world models (VWMs) synthesize interactive, action-conditioned rollouts from a single context image. However, it remains an open question how robust these models are to adve…
Q-ARVD: Quantizing Autoregressive Video Diffusion Models
Siao Tang, Xinyin Ma, Gongfan Fang +2
Autoregressive video diffusion models (ARVDs) have emerged as a promising architecture for streaming video generation, paving the way for real-time interactive video generation and…
ReactiveGWM: Steering NPC in Reactive Game World Models
Zeqing Wang, Danze Chen, Zhaohu Xing +4
Current game world models simulate environments from a subjective, player-centric perspective. However, by treating the Non-Player Character (NPC) merely as background pixels, thes…
Free Geometry: Refining 3D Reconstruction from Longer Versions of Itself
Yuhang Dai, Xingyi Yang
Feed-forward 3D reconstruction models are efficient but rigid: once trained, they perform inference in a zero-shot manner and cannot adapt to the test scene. As a result, visually…
WorldWarp: Propagating 3D Geometry with Asynchronous Video Diffusion
Hanyang Kong, Xingyi Yang, Xiaoxu Zheng +1
Generating long-range, geometrically consistent video presents a fundamental dilemma: while consistency demands strict adherence to 3D geometry in pixel space, state-of-the-art gen…