From the 2 of 58 linked papers with an AI index.
58 papers
SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling
Zheng Liu, Zijian He, Huiguo He +5
Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes, such as objects span differ…
Thermal Hall tomography of chiral superconductivity in rhombohedral graphene
Kumar Ghosh
A chiral superconductor carries chiral Majorana modes along its edges, and a single integer, the Bogoliubov--de Gennes Chern number, counts them. Thirty years of candidate material…
SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning
Keyang Zhong, Kuo Wang, Peng Liu +5
Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited co…
MoRoute: Dynamic Routing for In-Context Multimodal Video Generation
Chong Gao, Jie Ma, Zhan Peng +5
Multimodal video generation aims to generate and edit videos conditioned on arbitrary combinations of text, images, and videos within a single model, allowing diverse tasks to shar…
AdaAnchor4D: Anchor-Conditioned Spatiotemporal Feature Aggregation for Monocular UAV 4D Reconstruction
Peiyi Xu, Junpeng Zhang, Guanbin Li +6
The paper introduces AdaAnchor4D, a method that uses adaptive anchor-based feature aggregation to improve monocular UAV video reconstruction of dynamic urban scenes, reducing artif…
ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning
Zijie Wang, Wei Zhang, Weiming Zhang +4
ARDepth proposes an auto-regressive approach to monocular depth estimation that builds depth maps progressively across increasing spatial resolutions, using scale‑progressive condi…