2 papers
cs.CV2026
DepthART: Scaling Foundation Monocular Depth to Tiny Models
Feng Xue, Wu Chen, Mingshuai Zhao +7
Recent geometric foundation models (e.g., Metric3D, Depth Anything and UniDepth) have substantially improved monocular depth estimation (MDE) in both cross-scene generalization and…
cs.CV2026
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
Shida Gao, Feng Xue, Xiangfeng Wang +8
Multimodal large language models (MLLMs) are rapidly expanding from general video understanding to finer-grained understanding such as spatio-temporal video grounding (STVG) and re…