5 papers
Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
Junhao Xiao, Zhiyu Wu, Hao Lin +5
Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing me…
Path-Decoupled Hyperbolic Flow Matching for Few-Shot Adaptation
Lin Li, Ziqi Jiang, Gefan Ye +5
Recent advances in cross-modal few-shot adaptation treat visual-semantic alignment as a continuous feature transport problem via Flow Matching (FM). However, we argue that Euclidea…
Compositional Feature Augmentation for Unbiased Scene Graph Generation
Lin Li, Guikun Chen, Jun Xiao +3
Scene Graph Generation (SGG) aims to detect all the visual relation triplets \texttt{sub}, \texttt{pred}, \texttt{obj} in a given image. With the emergence of various advance…
Compositional Zero-shot Learning via Progressive Language-based Observations
Lin Li, Guikun Chen, Zhen Wang +2
Compositional zero-shot learning aims to recognize unseen state-object compositions by leveraging known primitives (state and object) during training. However, effectively modeling…
Decomposed Prototype Learning for Few-Shot Scene Graph Generation
Xingchen Li, Jun Xiao, Guikun Chen +4
Today's scene graph generation (SGG) models typically require abundant manual annotations to learn new predicate types. Therefore, it is difficult to apply them to real-world appli…