From the 1 of 4 linked papers with an AI index.
4 papers
ToxScreen: Detecting Whether an LLM Has Been Poisoned
Anthony Hughes, Nicole Xing, Collin Francel +2
The paper introduces ToxScreen, a benchmark of backdoored large language models, and evaluates methods for recovering hidden triggers under realistic defender constraints, finding…
Boundary-targeted Membership Inference Attacks on Safety Classifiers
Anthony Hughes, Alexander Goldberg, Prince Jha +3
Safety classifiers are essential safeguards within generative AI systems, filtering harmful content or identifying at-risk users when interacting with large language models. Despit…
PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing
Anthony Hughes, Vasisht Duddu, N. Asokan +2
Language models (LMs) may memorize personally identifiable information (PII) from training data, enabling adversaries to extract it during inference. Existing defense mechanisms su…
How Private are Language Models in Abstractive Summarization?
Anthony Hughes, Ning Ma, Nikolaos Aletras
In sensitive domains such as medical and legal, protecting sensitive information is critical, with protective laws strictly prohibiting the disclosure of personal data. This poses…