Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Defending Jailbreak Prompts via In-Context Adversarial Game
Yujun Zhou, Yufei Han, Haomin Zhuang +4
Large Language Models (LLMs) demonstrate remarkable capabilities across diverse applications. However, concerns regarding their security, particularly the vulnerability to jailbrea…
cs.LG2024
Manipulating Predictions over Discrete Inputs in Machine Teaching
Xiaodong Wu, Yufei Han, Hayssam Dahrouj +3
Machine teaching often involves the creation of an optimal (typically minimal) dataset to help a model (referred to as the `student') achieve specific goals given by a teacher. Whi…