activity
20242026
collaborators

21 papers

cs.CV2026

Visual Geometry Transformer in the Wild: Distractor-Free 3D Reconstruction

Tianbo Pan, Xingyi Yang, Shizun Wang +1

Current end-to-end multi-view 3D reconstruction methods achieve impressive results, but rely on a restrictive static assumption: the scenes is entire distractor-free with perfect c…

cs.LG2026

SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior

Mingyue Cui, Linghui Shen, Xingyi Yang

Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features. Recent latent-space defenses increasingly rely on these decompositions, assuming that…

cs.CV2026

BadWorld: Adversarial Attacks on World Models

Linghui Shen, Mingyue Cui, Xingyi Yang

Visual world models (VWMs) synthesize interactive, action-conditioned rollouts from a single context image. However, it remains an open question how robust these models are to adve…

cs.CV2026

Q-ARVD: Quantizing Autoregressive Video Diffusion Models

Siao Tang, Xinyin Ma, Gongfan Fang +2

Autoregressive video diffusion models (ARVDs) have emerged as a promising architecture for streaming video generation, paving the way for real-time interactive video generation and…

cs.CV2026

ReactiveGWM: Steering NPC in Reactive Game World Models

Zeqing Wang, Danze Chen, Zhaohu Xing +4

Current game world models simulate environments from a subjective, player-centric perspective. However, by treating the Non-Player Character (NPC) merely as background pixels, thes…

cs.CV2026

Free Geometry: Refining 3D Reconstruction from Longer Versions of Itself

Yuhang Dai, Xingyi Yang

Feed-forward 3D reconstruction models are efficient but rigid: once trained, they perform inference in a zero-shot manner and cannot adapt to the test scene. As a result, visually…