activity
20242026
most citedDepth Anything 3: Recovering the Visual Space from Any Views

2 citations · 2 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2026

Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning

Jian Shi, Michael Birsak, Wenqing Cui +2

This paper revisits the role of positional embeddings (PEs) within vision transformers (ViTs) from a geometric perspective. We show that PEs are not mere token indices but effectiv…

cs.CV20252 cited

Depth Anything 3: Recovering the Visual Space from Any Views

Haotong Lin, Sili Chen, Junhao Liew +5

We present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of…

cs.CV2025

PatchRefiner V2: Fast and Lightweight Real-Domain High-Resolution Metric Depth Estimation

Zhenyu Li, Wenqing Cui, Shariq Farooq Bhat +1

While current high-resolution depth estimation methods achieve strong results, they often suffer from computational inefficiencies due to reliance on heavyweight models and multipl…

cs.CV2024

Amodal Depth Anything: Amodal Depth Estimation in the Wild

Zhenyu Li, Mykola Lavreniuk, Jian Shi +2

Amodal depth estimation aims to predict the depth of occluded (invisible) parts of objects in a scene. This task addresses the question of whether models can effectively perceive t…

cs.CV2024

ImmersePro: End-to-End Stereo Video Synthesis Via Implicit Disparity Learning

Jian Shi, Zhenyu Li, Peter Wonka

We introduce \textit{ImmersePro}, an innovative framework specifically designed to transform single-view videos into stereo videos. This framework utilizes a novel dual-branch arch…