5 papers
DepthART: Scaling Foundation Monocular Depth to Tiny Models
Feng Xue, Wu Chen, Mingshuai Zhao +7
Recent geometric foundation models (e.g., Metric3D, Depth Anything and UniDepth) have substantially improved monocular depth estimation (MDE) in both cross-scene generalization and…
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
Shida Gao, Feng Xue, Xiangfeng Wang +8
Multimodal large language models (MLLMs) are rapidly expanding from general video understanding to finer-grained understanding such as spatio-temporal video grounding (STVG) and re…
NTIRE 2025 Challenge on HR Depth from Images of Specular and Transparent Surfaces
Pierluigi Zama Ramirez, Fabio Tosi, Luigi Di Stefano +36
This paper reports on the NTIRE 2025 challenge on HR Depth From images of Specular and Transparent surfaces, held in conjunction with the New Trends in Image Restoration and Enhanc…
Cues3D: Unleashing the Power of Sole NeRF for Consistent and Unique Instances in Open-Vocabulary 3D Panoptic Segmentation
Feng Xue, Wenzhuang Xu, Guofeng Zhong +2
Open-vocabulary 3D panoptic segmentation has recently emerged as a significant trend. Top-performing methods currently integrate 2D segmentation with geometry-aware 3D primitives.…
The Fourth Monocular Depth Estimation Challenge
Anton Obukhov, Matteo Poggi, Fabio Tosi +54
This paper presents the results of the fourth edition of the Monocular Depth Estimation Challenge (MDEC), which focuses on zero-shot generalization to the SYNS-Patches benchmark, a…