5 papers
How Context Attribution Handles What the Model Already Knows
Quoc-Huy Trinh, Lin Zhu, Sebastian Szyller
Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works show the initial success in attributing th…
Amulet: a Python Library for Assessing Interactions Among ML Defenses and Risks
Asim Waheed, Vasisht Duddu, Rui Zhang +1
Machine learning (ML) models are susceptible to various risks to security, privacy, and fairness. Most defenses are designed to protect against each risk individually (intended int…
Soft Token Attacks Cannot Reliably Audit Unlearning in Large Language Models
Haokun Chen, Sebastian Szyller, Weilin Xu +1
Large language models (LLMs) are trained using massive datasets, which often contain undesirable content such as harmful texts, personal information, and copyrighted material. To a…
Atlas: A Framework for ML Lifecycle Provenance & Transparency
Marcin Spoczynski, Marcela S. Melara, Sebastian Szyller
The rapid adoption of open source machine learning (ML) datasets and models exposes today's AI applications to critical risks like data poisoning and supply chain attacks across th…
Imperceptible Adversarial Examples in the Physical World
Weilin Xu, Sebastian Szyller, Cory Cornelius +5
Adversarial examples in the digital domain against deep learning-based computer vision models allow for perturbations that are imperceptible to human eyes. However, producing simil…