6 papers
Efficiency Follows Global-Local Decoupling
Zhenyu Yang, Gensheng Pei, Tao Chen +4
Modern vision models must capture image-level context without sacrificing local detail while remaining computationally affordable. We revisit this tradeoff and advance a simple pri…
Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth Estimation
Xinhao Cai, Gensheng Pei, Zeren Sun +3
In this paper, we propose \textbf{Iris}, a deterministic framework for Monocular Depth Estimation (MDE) that integrates real-world priors into the diffusion model. Conventional fee…
PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation
Jianjian Yin, Tao Chen, Yi Chen +4
Recent advances in vision-language models (VLMs) have garnered substantial attention in open-vocabulary semantic and part segmentation (OSPS). However, existing methods extract ima…
Towards Remote Sensing Change Detection with Neural Memory
Zhenyu Yang, Gensheng Pei, Yazhou Yao +3
Remote sensing change detection is essential for environmental monitoring, urban planning, and related applications. However, current methods often struggle to capture long-range d…
Taming SAM3 in the Wild: A Concept Bank for Open-Vocabulary Segmentation
Gensheng Pei, Xiruo Jiang, Yazhou Yao +3
The recent introduction of \texttt{SAM3} has revolutionized Open-Vocabulary Segmentation (OVS) through \textit{promptable concept segmentation}, which grounds pixel predictions in…
Combating Noisy Labels through Fostering Self- and Neighbor-Consistency
Zeren Sun, Yazhou Yao, Tongliang Liu +3
Label noise is pervasive in various real-world scenarios, posing challenges in supervised deep learning. Deep networks are vulnerable to such label-corrupted samples due to the mem…