3 citations · 5 across the 3 of their papers we have counts for
6 papers · 1 filter
Local Margin Restoration for Test-Time Adaptation of Vision-Language Models
Yan Huang, Guowei Wang, Xu Wang +2
Vision-language models (VLMs) such as CLIP exhibit remarkable zero-shot capabilities, yet their performance frequently degrades sharply under unexpected test-time distribution shif…
SGC-Net: Stratified Granular Comparison Network for Open-Vocabulary HOI Detection
Xin Lin, Chong Shi, Zuopeng Yang +2
Recent open-vocabulary human-object interaction (OV-HOI) detection methods primarily rely on large language model (LLM) for generating auxiliary descriptions and leverage knowledge…
Distraction is All You Need for Multimodal Large Language Model Jailbreaking
Zuopeng Yang, Jiluan Fan, Anli Yan +5
Multimodal Large Language Models (MLLMs) bridge the gap between visual and textual data, enabling a range of advanced applications. However, complex internal interactions among vis…
HL-Net: Heterophily Learning Network for Scene Graph Generation
Xin Lin, Changxing Ding, Yibing Zhan +2
Scene graph generation (SGG) aims to detect objects and predict their pairwise relationships within an image. Current SGG methods typically utilize graph neural networks (GNNs) to…
RU-Net: Regularized Unrolling Network for Scene Graph Generation
Xin Lin, Changxing Ding, Jing Zhang +2
Scene graph generation (SGG) aims to detect objects and predict the relationships between each pair of objects. Existing SGG methods usually suffer from several issues, including 1…
GPS-Net: Graph Property Sensing Network for Scene Graph Generation
Xin Lin, Changxing Ding, Jinquan Zeng +1
Scene graph generation (SGG) aims to detect objects in an image along with their pairwise relationships. There are three key properties of scene graph that have been underexplored…