6 papers
On the Adversarial Robustness of Multimodal LLM Judges
Zihan Wang, Guansong Pang, Zelin Liu +3
Multimodal Large Language Models (MLLMs) are increasingly used as automated judges, e.g., for image quality and safety assessment. However, their adversarial robustness remains lar…
TokenCLIP: Token-wise Prompt Learning for Zero-shot Anomaly Detection
Qihang Zhou, Binbin Gao, Guansong Pang +3
Adapting CLIP for anomaly detection on unseen objects has shown strong potential in a zero-shot manner. However, existing methods typically rely on a single textual space to align…
AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds
Qizhou Wang, Hanxun Huang, Guansong Pang +2
Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, th…
MTAttack: Multi-Target Backdoor Attacks against Large Vision-Language Models
Zihan Wang, Guansong Pang, Wenjun Miao +2
Recent advances in Large Visual Language Models (LVLMs) have demonstrated impressive performance across various vision-language tasks by leveraging large-scale image-text pretraini…
SEMPO: Lightweight Foundation Models for Time Series Forecasting
Hui He, Kun Yi, Yuanchi Ma +3
The recent boom of large pre-trained models witnesses remarkable success in developing foundation models (FMs) for time series forecasting. Despite impressive performance across di…
DOEI: Dual Optimization of Embedding Information for Attention-Enhanced Class Activation Maps
Hongjie Zhu, Zeyu Zhang, Guansong Pang +6
Weakly supervised semantic segmentation (WSSS) typically utilizes limited semantic annotations to obtain initial Class Activation Maps (CAMs). However, due to the inadequate coupli…