5 papers
ICED: Concept-level Machine Unlearning via Interpretable Concept Decomposition
Shen Lin, Jing Lin, Junhao Dong +2
Machine unlearning in Vision-Language Models (VLMs) is typically performed at the image or instance level, making it difficult to precisely remove target knowledge without affectin…
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion
ShiYing Huang, Liang Lin, Yuer Li +6
In the realm of multi-objective alignment for large language models, balancing disparate human preferences often manifests as a zero-sum conflict. Specifically, the intrinsic tensi…
GaitProtector: Impersonation-Driven Gait De-Identification via Training-Free Diffusion Latent Optimization
Huiran Duan, Qian Zhou, Zhongliang Guo +4
Conventional gait de-identification methods often encounter an inherent trade-off: they either provide insufficient identity suppression or introduce spatiotemporal distortions tha…
BiAxisBias: Evaluating LLM Bias Beyond a Single Prompt and a Single Explanation
Jialing Gan, Junhao Dong, Songze Li
LLM bias scores can depend on audit design. We introduce BiAxisBias, a prespecified audit varying task, role, perspective, sentiment, and wording over 200 stereotype statements whi…
A Gray-box Attack against Latent Diffusion Model-based Image Editing by Posterior Collapse
Zhongliang Guo, Chun Tong Lei, Lei Fang +7
Recent advancements in Latent Diffusion Models (LDMs) have revolutionized image synthesis and manipulation, raising significant concerns about data misappropriation and intellectua…