6 papers
Effective Multi-sensor Conditioning for Street-view Novel-view Synthesis
Zhengfei Kuang, Adam Sun, Liyuan Zhu +7
Modern vehicle platforms are equipped with a rich sensor suite, including LiDAR, calibrated multi-camera rigs, and accurate ego-motion, that in principle offers strong signal for r…
GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation
Jan Ackermann, Shengqu Cai, Boyang Deng +3
Generating geometrically consistent videos remains an open challenge: text-to-video diffusion models trained on web-scale data treat geometry only implicitly, leading to object def…
VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
Zhengfei Kuang, Rui Lin, Long Zhao +3
Despite the remarkable progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, their application to complex 3D scene manipulation remains underexplored. I…
Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion
Yuming Gu, Yizhi Wang, Yining Hong +9
Embodied visual planning aims to enable manipulation tasks by imagining how a scene evolves toward a desired goal and using the imagined trajectories to guide actions. Video diffus…
BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
Yiming Wang, Qihang Zhang, Shengqu Cai +7
Emerging video diffusion models achieve high visual fidelity but fundamentally couple scene dynamics with camera motion, limiting their ability to provide precise spatial and tempo…
X-Dyna: Expressive Dynamic Human Image Animation
Di Chang, Hongyi Xu, You Xie +12
We introduce X-Dyna, a novel zero-shot, diffusion-based pipeline for animating a single human image using facial expressions and body movements derived from a driving video, that g…