12 papers
PEARL: Geometry Aligns Semantics for Training-Free Open-Vocabulary Semantic Segmentation
Gensheng Pei, Xiruo Jiang, Xinhao Cai +3
Training-free open-vocabulary semantic segmentation (OVSS) promises rapid adaptation to new label sets without retraining. Yet, many methods rely on heavy post-processing or handle…
Efficiency Follows Global-Local Decoupling
Zhenyu Yang, Gensheng Pei, Tao Chen +4
Modern vision models must capture image-level context without sacrificing local detail while remaining computationally affordable. We revisit this tradeoff and advance a simple pri…
Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth Estimation
Xinhao Cai, Gensheng Pei, Zeren Sun +3
In this paper, we propose \textbf{Iris}, a deterministic framework for Monocular Depth Estimation (MDE) that integrates real-world priors into the diffusion model. Conventional fee…
PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation
Jianjian Yin, Tao Chen, Yi Chen +4
Recent advances in vision-language models (VLMs) have garnered substantial attention in open-vocabulary semantic and part segmentation (OSPS). However, existing methods extract ima…
PKINet-v2: Towards Powerful and Efficient Poly-Kernel Remote Sensing Object Detection
Xinhao Cai, Liulei Li, Gensheng Pei +3
Object detection in remote sensing images (RSIs) is challenged by the coexistence of geometric and spatial complexity: targets may appear with diverse aspect ratios, while spanning…
Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis
Xinhao Cai, Liulei Li, Gensheng Pei +4
This paper presents a generation-based debiasing framework for object detection. Prior debiasing methods are often limited by the representation diversity of samples, while naive g…