4 papers
DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video Generation
Jian Shi, Qian Wang, Zhenyu Li +3
Generating high-quality stereo videos requires consistent depth perception and temporal coherence across frames. Despite advances in image and video synthesis using diffusion model…
Any Resolution Any Geometry: From Multi-View To Multi-Patch
Wenqing Cui, Zhenyu Li, Mykola Lavreniuk +4
Joint estimation of surface normals and depth is essential for holistic 3D scene understanding, yet high-resolution prediction remains difficult due to the trade-off between preser…
Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning
Jian Shi, Michael Birsak, Wenqing Cui +2
This paper revisits the role of positional embeddings (PEs) within vision transformers (ViTs) from a geometric perspective. We show that PEs are not mere token indices but effectiv…
Amodal Depth Anything: Amodal Depth Estimation in the Wild
Zhenyu Li, Mykola Lavreniuk, Jian Shi +2
Amodal depth estimation aims to predict the depth of occluded (invisible) parts of objects in a scene. This task addresses the question of whether models can effectively perceive t…