most citedFine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

41 citations · 58 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CV2024

JIGMARK: A Black-Box Approach for Enhancing Image Watermarks against Diffusion Model Edits

Minzhou Pan, Yi Zeng, Xue Lin +4

In this study, we investigate the vulnerability of image watermarks to diffusion-model-based image editing, a challenge exacerbated by the computational cost of accessing gradient…

cs.AI202411 cited

A Safe Harbor for AI Evaluation and Red Teaming

Shayne Longpre, Sayash Kapoor, Kevin Klyman +20

Independent evaluation and red teaming are critical for identifying the risks posed by generative AI systems. However, the terms of service and enforcement strategies used by promi…

cs.CR2024

RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content

Zhuowen Yuan, Zidi Xiong, Yi Zeng +4

Recent advancements in Large Language Models (LLMs) have showcased remarkable capabilities across various tasks in different domains. However, the emergence of biases and the poten…

cs.CL20246 cited

How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Yi Zeng, Hongpeng Lin, Jingwen Zhang +3

Most traditional AI safety research has approached AI models as machines and centered on algorithm-focused attacks developed by security experts. As large language models (LLMs) be…

cs.CR2023

Who Leaked the Model? Tracking IP Infringers in Accountable Federated Learning

Shuyang Yu, Junyuan Hong, Yi Zeng +3

Federated learning (FL) emerges as an effective collaborative learning framework to coordinate data and computation resources from massive and distributed clients in training. Such…

cs.CL202341 cited

Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Xiangyu Qi, Yi Zeng, Tinghao Xie +4

Optimizing large language models (LLMs) for downstream use cases often involves the customization of pre-trained LLMs through further fine-tuning. Meta's open release of Llama mode…