1 paper
Lihua Zhou, Mao Ye, Xiatian Zhu +7
Open-vocabulary object detection with vision-language models (VLMs) such as Grounding DINO suffers from performance degradation under test-time distribution shifts, primarily due t…