3 citations · 3 across the 4 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.LG2024★ 3 cited
Defending Jailbreak Prompts via In-Context Adversarial Game
Yujun Zhou, Yufei Han, Haomin Zhuang +4
Large Language Models (LLMs) demonstrate remarkable capabilities across diverse applications. However, concerns regarding their security, particularly the vulnerability to jailbrea…
cs.LG2024
Manipulating Predictions over Discrete Inputs in Machine Teaching
Xiaodong Wu, Yufei Han, Hayssam Dahrouj +3
Machine teaching often involves the creation of an optimal (typically minimal) dataset to help a model (referred to as the `student') achieve specific goals given by a teacher. Whi…