4 papers
Comment and Control: Hijacking Agentic Workflows via Context-Grounded Evolution
Neil Fendley, Zhengyu Liu, Aonan Guan +2
Automation platforms such as GitHub Actions and n8n are increasingly adopting so-called agentic workflows, which integrate Large Language Model (LLM) agents for tasks such as code…
Trojans in Artificial Intelligence (TrojAI) Final Report
Kristopher W. Reese, Taylor Kulp-McDowall, Michael Majurski +68
The Intelligence Advanced Research Projects Activity (IARPA) launched the TrojAI program to confront an emerging vulnerability in modern artificial intelligence: the threat of AI T…
A Systematic Review of Poisoning Attacks Against Large Language Models
Neil Fendley, Edward W. Staley, Joshua Carney +3
With the widespread availability of pretrained Large Language Models (LLMs) and their training datasets, concerns about the security risks associated with their usage has increased…
Mirror Mirror on the Wall, Have I Forgotten it All? A New Framework for Evaluating Machine Unlearning
Brennon Brimhall, Philip Mathew, Neil Fendley +2
Machine unlearning methods take a model trained on a dataset and a forget set, then attempt to produce a model as if it had only been trained on the examples not in the forget set.…