7 papers
QuerySplat: Decoupling Geometry and Appearance Representations in 3DGS Prediction
Yinglong Li, Donghui Shen, Xiaoyu Zhang +5
While feed-forward 3D Gaussian Splatting (3DGS) enables efficient 3D reconstruction, achieving high-fidelity rendering remains challenging. Existing pixel-aligned approaches suffer…
PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention
Yipeng Chen, Zhichao Ye, Zhenzhou Fang +6
We propose PostCam, a streamlined framework for novel-view video generation that achieves superior detail preservation and precise camera trajectory editing in dynamic scenes. Curr…
SEMAGIC: Learning Semantically Consistent Deformable 3D Representations from In-the-Wild Images
Sky Cen, Wufei Ma, Guofeng Zhang +2
Learning deformable 3D object models from single-view in-the-wild images has enabled impressive 3D shape reconstruction without supervision. However, it remains unclear whether the…
InSpatio-WorldFM: An Open-Source Real-Time Generative Frame Model
InSpatio Team, Donghui Shen, Guofeng Zhang +16
We present InSpatio-WorldFM, an open-source real-time frame model for spatial intelligence. Unlike video-based world models that rely on sequential frame generation and incur subst…
INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling
InSpatio Team, Donghui Shen, Guofeng Zhang +20
Building world models with spatial consistency and real-time interactivity remains a fundamental challenge in computer vision. Current video generation paradigms often struggle wit…
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation
Shangjin Zhai, Zhichao Ye, Jialin Liu +10
Recent advances in large reconstruction and generative models have significantly improved scene reconstruction and novel view generation. However, due to compute limitations, each…