7 papers
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
Zhenxing Zhang, Yaxiong Wang, Lechao Cheng +3
We present ASAP, a new framework for detecting and grounding multi-modal media manipulation (DGM4).Upon thorough examination, we observe that accurate fine-grained cross-modal sema…
SSAM: Self-Supervised Association Modeling for Test-Time Adaption
Yaxiong Wang, Zhenqiang Zhang, Lechao Cheng +3
Test-time adaption (TTA) has witnessed important progress in recent years, the prevailing methods typically first encode the image and the text and design strategies to model the a…
Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training Approach
Yan Zhang, Lechao Cheng, Yaxiong Wang +2
Micro-Action Recognition (MAR) aims to classify subtle human actions in video. However, annotating MAR datasets is particularly challenging due to the subtlety of actions. To this…
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
Yaxiong Wang, Yujiao Wu, Lianwei Wu +3
Recent advancements in image-text matching have been notable, yet prevailing models predominantly cater to broad queries and struggle with accommodating fine-grained query intentio…
Knowledge Swapping via Learning and Unlearning
Mingyu Xing, Lechao Cheng, Shengeng Tang +3
We introduce \textbf{Knowledge Swapping}, a novel task designed to selectively regulate knowledge of a pretrained model by enabling the forgetting of user\-specified information, r…
Navigating Semantic Drift in Task-Agnostic Class-Incremental Learning
Fangwen Wu, Lechao Cheng, Shengeng Tang +4
Class-incremental learning (CIL) seeks to enable a model to sequentially learn new classes while retaining knowledge of previously learned ones. Balancing flexibility and stability…