12 papers
Reward-Guided Semantic Evolution for Test-time Adaptive Object Detection
Lihua Zhou, Mao Ye, Xiatian Zhu +7
Open-vocabulary object detection with vision-language models (VLMs) such as Grounding DINO suffers from performance degradation under test-time distribution shifts, primarily due t…
FACTOR: Counterfactual Training-Free Test-Time Adaptation for Open-Vocabulary Object Detection
Kaixiang Zhao, Mao Ye, Lihua Zhou +4
Open-vocabulary object detection often fails under distribution shifts, as it can be misled by spurious correlations between non-causal visual attributes (e.g., brightness, texture…
Source-Free Domain Adaptation with Vision-Language Prior
Song Tang, Yunxiang Bai, Wenxin Su +3
Source-Free Domain Adaptation (SFDA) seeks to adapt a source model, which is pre-trained on a supervised source domain, for a target domain, with only access to unlabeled target tr…
Consistent text-to-image generation via scene de-contextualization
Song Tang, Peihao Gong, Kunyu Li +5
Consistent text-to-image (T2I) generation seeks to produce identity-preserving images of the same subject across diverse scenes, yet it often fails due to a phenomenon called ident…
Unified Source-Free Domain Adaptation
Song Tang, Wenxin Su, Mao Ye +2
In the pursuit of transferring a source model to a target domain without access to the source training data, Source-Free Domain Adaptation (SFDA) has been extensively explored acro…
Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
Lihua Zhou, Mao Ye, Shuaifeng Li +7
Vision-language models (VLMs) such as CLIP and Grounding DINO have achieved remarkable success in object recognition and detection. However, their performance often degrades under…