works on

From the 1 of 12 linked papers with an AI index.

collaborators

12 papers

cs.CV2026

Wonder: Video World Model Done Better

Jiacong Xu, Hanwen Jiang, Zhixin Shu +3

Wonder is a video world model that lets users explore a generated scene in real time by moving a virtual camera, using a dense coordinate conditioning and a sparse attention memory…

cs.CV2026

SPEAR: A Simulator for Photorealistic Embodied AI Research

Mike Roberts, Renhan Wang, Rushikesh Zawar +10

Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited gene…

cs.CV2026

LooseControlVideo: Directorial Video Control using Spatial Blocking

Shariq Farooq Bhat, Niloy J. Mitra, Kalyan Sunkavalli

Precise 3D spatial orchestration in text-to-video generation remains a significant challenge, particularly for multi-object scenes where semantic layout and temporal dynamics are o…

cs.CV2026

OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation

Yuheng Liu, Xin Lin, Xinke Li +9

Modeling scenes using video generation models has garnered growing research interest in recent years. However, most existing approaches rely on perspective video models that synthe…

cs.CV2026

E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training

Qitao Zhao, Hao Tan, Qianqian Wang +5

Self-supervised pre-training has driven rapid progress in foundation models for language, 2D images, and video, yet remains largely unexplored for learning 3D-aware representations…

cs.CV2026

tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction

Chen Wang, Hao Tan, Wang Yifan +6

We propose tttLRM, a novel large 3D reconstruction model that leverages a Test-Time Training (TTT) layer to enable long-context, autoregressive 3D reconstruction with linear comput…