4 papers · 1 filter
Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models
Sultan Alshehri, Zhantao Yang, Han Zhang +1
Dual-encoder vision-language models (VLMs) expose a similarity interface that enables zero-shot retrieval but fails compositional constraints: queries like "umbrella and no person"…
A Reference-Based 3D Semantic-Aware Framework for Accurate Local Facial Attribute Editing
Yu-Kai Huang, Yutong Zheng, Yen-Shuo Su +4
Facial attribute editing plays a crucial role in synthesizing realistic faces with specific characteristics while maintaining realistic appearances. Despite advancements, challenge…
RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection
Fangyi Chen, Han Zhang, Zhantao Yang +3
Open-vocabulary object detection (OVD) requires solid modeling of the region-semantic relationship, which could be learned from massive region-text pairs. However, such data is lim…
Solving Missing-Annotation Object Detection with Background Recalibration Loss
Han Zhang, Fangyi Chen, Zhiqiang Shen +3
This paper focuses on a novel and challenging detection scenario: A majority of true objects/instances is unlabeled in the datasets, so these missing-labeled areas will be regarded…