From the 2 of 64 linked papers with an AI index.
65 papers · 1 filter
SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling
Zheng Liu, Zijian He, Huiguo He +5
Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes, such as objects span differ…
MoRoute: Dynamic Routing for In-Context Multimodal Video Generation
Chong Gao, Jie Ma, Zhan Peng +5
Multimodal video generation aims to generate and edit videos conditioned on arbitrary combinations of text, images, and videos within a single model, allowing diverse tasks to shar…
AdaAnchor4D: Anchor-Conditioned Spatiotemporal Feature Aggregation for Monocular UAV 4D Reconstruction
Peiyi Xu, Junpeng Zhang, Guanbin Li +6
The paper introduces AdaAnchor4D, a method that uses adaptive anchor-based feature aggregation to improve monocular UAV video reconstruction of dynamic urban scenes, reducing artif…
ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning
Zijie Wang, Wei Zhang, Weiming Zhang +4
ARDepth proposes an auto-regressive approach to monocular depth estimation that builds depth maps progressively across increasing spatial resolutions, using scale‑progressive condi…
DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior
Junjia Huang, Binbin Yang, Pengxiang Yan +6
Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scen…
PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation
Weixing Chen, Zhuoqian Feng, Yang Liu +6
Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense objec…