15 citations · 45 across the 18 of their papers we have counts for
18 papers
Segment Anything for Videos: A Systematic Survey
Chunhui Zhang, Yawen Cui, Weilin Lin +4
The recent wave of foundation models has witnessed tremendous success in computer vision (CV) and beyond, with the segment anything model (SAM) having sparked a passion for explori…
T2IShield: Defending Against Backdoors on Text-to-Image Diffusion Models
Zhongqi Wang, Jie Zhang, Shiguang Shan +1
While text-to-image diffusion models demonstrate impressive generation capabilities, they also exhibit vulnerability to backdoor attacks, which involve the manipulation of model ou…
HPNet: Dynamic Trajectory Forecasting with Historical Prediction Attention
Xiaolong Tang, Meina Kan, Shiguang Shan +3
Predicting the trajectories of road agents is essential for autonomous driving systems. The recent mainstream methods follow a static paradigm, which predicts the future trajectory…
GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
Hao Lu, Xuesong Niu, Jiyao Wang +12
Multimodal large language models (MLLMs) are designed to process and integrate information from multiple sources, such as text, speech, images, and videos. Despite its success in l…
Task Attribute Distance for Few-Shot Learning: Theoretical Analysis and Applications
Minyang Hu, Hong Chang, Zong Guo +3
Few-shot learning (FSL) aims to learn novel tasks with very few labeled samples by leveraging experience from \emph{related} training tasks. In this paper, we try to understand FSL…
Contrastive Learning of Person-independent Representations for Facial Action Unit Detection
Yong Li, Shiguang Shan
Facial action unit (AU) detection, aiming to classify AU present in the facial image, has long suffered from insufficient AU annotations. In this paper, we aim to mitigate this dat…