5 papers · 1 filter
Semantic-aware Adversarial Fine-tuning for CLIP
Jiacheng Zhang, Jinhao Li, Hanxun Huang +3
Recent studies have shown that CLIP model's adversarial robustness in zero-shot classification tasks can be enhanced by adversarially fine-tuning its image encoder with adversarial…
Exploring Weak-to-Strong Generalization for CLIP-based Classification
Jinhao Li, Sarah M. Erfani, Lei Feng +2
Aligning large-scale commercial models with user intent is crucial to preventing harmful outputs. Current methods rely on human supervision but become impractical as model complexi…
Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
Xueqi Ma, Yanbei Jiang, Sarah Erfani +4
Multimodal Large Language Models (MLLMs) have demonstrated exceptional performance across various objective multimodal perception tasks, yet their application to subjective, emotio…
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
Hanxun Huang, Sarah Erfani, Yige Li +2
As Contrastive Language-Image Pre-training (CLIP) models are increasingly adopted for diverse downstream tasks and integrated into large vision-language models (VLMs), their suscep…
Efficient Neural Implicit Representation for 3D Human Reconstruction
Zexu Huang, Sarah Monazam Erfani, Siying Lu +1
High-fidelity digital human representations are increasingly in demand in the digital world, particularly for interactive telepresence, AR/VR, 3D graphics, and the rapidly evolving…