6 papers
MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold
Yang Zhou, Ziheng Wang, Yuqin Lu +4
We present MoVerse, a real-time video world model that creates an interactively navigable scene from a single narrow-field-of-view image. This setting is challenging because the in…
MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics
Haofeng Liu, Yang Zhou, Ziheng Wang +6
Generative novel view synthesis faces a fundamental dilemma: geometric priors provide spatial alignment but become sparse and inaccurate under view changes, while appearance priors…
Self-Corrected Image Generation with Explainable Latent Rewards
Yinyi Luo, Hrishikesh Gokhale, Marios Savvides +2
Despite significant progress in text-to-image generation, aligning outputs with complex prompts remains challenging, particularly for fine-grained semantics and spatial relations.…
Teacher-Student Diffusion Model for Text-Driven 3D Hand Motion Generation
Ching-Lam Cheng, Bin Zhu, Shengfeng He
Generating realistic 3D hand motion from natural language is vital for VR, robotics, and human-computer interaction. Existing methods either focus on full-body motion, overlooking…
Gimbal360: Canonicalizing Planar Diffusion for Spherical Panorama Completion
Yuqin Lu, Haofeng Liu, Yang Zhou +5
Diffusion models provide powerful priors for 2D image completion, but these priors are learned on bounded planar images and do not transfer directly to panoramas. Persp…
Lagrangian Motion Fields for Long-term Motion Generation
Yifei Yang, Zikai Huang, Chenshu Xu +1
Long-term motion generation is a challenging task that requires producing coherent and realistic sequences over extended durations. Current methods primarily rely on framewise moti…