3 papers
cs.CV2026
DepthART: Scaling Foundation Monocular Depth to Tiny Models
Feng Xue, Wu Chen, Mingshuai Zhao +7
Recent geometric foundation models (e.g., Metric3D, Depth Anything and UniDepth) have substantially improved monocular depth estimation (MDE) in both cross-scene generalization and…
cs.CV2026
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
Shida Gao, Feng Xue, Xiangfeng Wang +8
Multimodal large language models (MLLMs) are rapidly expanding from general video understanding to finer-grained understanding such as spatio-temporal video grounding (STVG) and re…
cs.LG2025
Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models
Yifan Jia, Kailin Jiang, Yuyang Liang +11
Large Multimodal Models(LMMs) face notable challenges when encountering multimodal knowledge conflicts, particularly under retrieval-augmented generation(RAG) frameworks where the…