18 citations · 34 across the 8 of their papers we have counts for
16 papers · 1 filter
Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
Junhao Xiao, Zhiyu Wu, Hao Lin +5
Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing me…
Path-Decoupled Hyperbolic Flow Matching for Few-Shot Adaptation
Lin Li, Ziqi Jiang, Gefan Ye +5
Recent advances in cross-modal few-shot adaptation treat visual-semantic alignment as a continuous feature transport problem via Flow Matching (FM). However, we argue that Euclidea…
DECap: Towards Generalized Explicit Caption Editing via Diffusion Mechanism
Zhen Wang, Xinyun Jiang, Jun Xiao +2
Explicit Caption Editing (ECE) -- refining reference image captions through a sequence of explicit edit operations (e.g., KEEP, DETELE) -- has raised significant attention due to i…
Compositional Zero-shot Learning via Progressive Language-based Observations
Lin Li, Guikun Chen, Zhen Wang +2
Compositional zero-shot learning aims to recognize unseen state-object compositions by leveraging known primitives (state and object) during training. However, effectively modeling…
Compositional Feature Augmentation for Unbiased Scene Graph Generation
Lin Li, Guikun Chen, Jun Xiao +3
Scene Graph Generation (SGG) aims to detect all the visual relation triplets \texttt{sub}, \texttt{pred}, \texttt{obj} in a given image. With the emergence of various advance…
Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph Generation
Wenqing Wang, Kaifeng Gao, Yawei Luo +5
Video-based scene graph generation (VidSGG) is an approach that aims to represent video content in a dynamic graph by identifying visual entities and their relationships. Due to th…