32 citations · 99 across the 77 of their papers we have counts for
94 papers · 1 filter
SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling
Zheng Liu, Zijian He, Huiguo He +5
Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes, such as objects span differ…
MoRoute: Dynamic Routing for In-Context Multimodal Video Generation
Chong Gao, Jie Ma, Zhan Peng +5
Multimodal video generation aims to generate and edit videos conditioned on arbitrary combinations of text, images, and videos within a single model, allowing diverse tasks to shar…
AdaAnchor4D: Anchor-Conditioned Spatiotemporal Feature Aggregation for Monocular UAV 4D Reconstruction
Peiyi Xu, Junpeng Zhang, Guanbin Li +6
Monocular UAV videos provide valuable observations for dynamic reconstruction of complex urban scenes. However, such scenes exhibit pronounced spatiotemporal heterogeneity: differe…
ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning
Zijie Wang, Wei Zhang, Weiming Zhang +4
Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that depth can be recovered as a globally smooth f…
PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation
Weixing Chen, Zhuoqian Feng, Yang Liu +6
Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense objec…
Non-Forgetting Knowledge Allocation with Bi-level Competition for Class-Incremental Learning
Xiang Tan, Run He, Yawen Cui +6
Class-Incremental Learning (CIL) with pre-trained models (PTMs) aims to sequentially adapt PTMs to new categories without forgetting old knowledge. Built upon PTMs, existing adapte…