5 papers
FALCON: Fine-grained Activation Manipulation by Contrastive Orthogonal Unalignment for Large Language Model
Jinwei Hu, Zhenglin Huang, Xiangyu Yin +4
Large language models have been widely applied, but can inadvertently encode sensitive or harmful information, raising significant safety concerns. Machine unlearning has emerged t…
Trustworthy Text-to-Image Diffusion Models: A Timely and Focused Survey
Yi Zhang, Zhen Chen, Chih-Hong Cheng +6
Text-to-Image (T2I) Diffusion Models (DMs) have garnered widespread attention for their impressive advancements in image generation. However, their growing popularity has raised et…
A Black-Box Evaluation Framework for Semantic Robustness in Bird's Eye View Detection
Fu Wang, Yanghao Zhang, Xiangyu Yin +4
Camera-based Bird's Eye View (BEV) perception models receive increasing attention for their crucial role in autonomous driving, a domain where concerns about the robustness and rel…
ProTIP: Probabilistic Robustness Verification on Text-to-Image Diffusion Models against Stochastic Perturbation
Yi Zhang, Yun Tang, Wenjie Ruan +4
Text-to-Image (T2I) Diffusion Models (DMs) have shown impressive abilities in generating high-quality images based on simple text descriptions. However, as is common with many Deep…
Building Guardrails for Large Language Models
Yi Dong, Ronghui Mu, Gaojie Jin +6
As Large Language Models (LLMs) become more integrated into our daily lives, it is crucial to identify and mitigate their risks, especially when the risks can have profound impacts…