4 papers
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
Youze Wang, Wenbo Hu, Yinpeng Dong +3
Large Language Models (LLMs) have evolved into Multimodal Large Language Models (MLLMs), significantly enhancing their capabilities by integrating visual information and other type…
Aligned Contrastive Loss for Long-Tailed Recognition
Jiali Ma, Jiequan Cui, Maeno Kazuki +4
In this paper, we propose an Aligned Contrastive Learning (ACL) algorithm to address the long-tailed recognition problem. Our findings indicate that while multi-view training boost…
Learning to Animate Images from A Few Videos to Portray Delicate Human Actions
Haoxin Li, Yingchen Yu, Qilong Wu +3
Despite recent progress, video generative models still struggle to animate static images into videos that portray delicate human actions, particularly when handling uncommon or nov…
CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual Interaction
Yuan Zhou, Qingshan Xu, Jiequan Cui +4
Recently, large efforts have been made to design efficient linear-complexity visual Transformers. However, current linear attention models are generally unsuitable to be deployed i…