collaborators

10 papers

cs.CV2026

RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection

Zhihao Zhang, Gengwei Zhang, Tianlong Chen +1

Monocular 3D object detection spans two regimes: closed-set detectors operating within a fixed category vocabulary, and open-vocabulary detectors that localize arbitrary categories…

cs.CV2026

Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth

Jung-Hee Kim, Xiaoming Liu

Monocular depth foundation models have demonstrated remarkable generalization capabilities across diverse environments. However, they continue to struggle with metric depth estimat…

cs.CV2026

DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

Jie Zhu, Girish Chandar Ganesan, Xiaoming Liu

Monocular metric depth estimation has achieved strong progress with large-scale training and universal-camera modeling, yet robust deployment across diverse camera settings, such a…

cs.CV2026

H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning

Zhanbo Huang, Xiaoming Liu, Yu Kong

Parametric human models capture global pose but cannot represent the non-rigid surface dynamics of clothing and soft tissue. Generic scene flow estimates dense motion but breaks do…

cs.CV2026

UniDAC: Universal Metric Depth Estimation for Any Camera

Girish Chandar Ganesan, Yuliang Guo, Liu Ren +1

Monocular metric depth estimation (MMDE) is a core challenge in computer vision, playing a pivotal role in real-world applications that demand accurate spatial understanding. Altho…

cs.CV2026

Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection

Zhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan +1

Monocular 3D detection (Mono3D) aims to infer 3D bounding boxes from a single RGB image. Without auxiliary sensors such as LiDAR, this task is inherently ill-posed since the 3D-to-…