4 papers
Back to Blackwell: Closing the Loop on Intransitivity in Multi-Objective Preference Fine-Tuning
Jiahao Zhang, Lujing Zhang, Keltin Grimes +3
A recurring challenge in preference fine-tuning (PFT) is handling (i.e., cyclic) preferences. Intransitive preferences often stem from either …
Trojans in Artificial Intelligence (TrojAI) Final Report
Kristopher W. Reese, Taylor Kulp-McDowall, Michael Majurski +68
The Intelligence Advanced Research Projects Activity (IARPA) launched the TrojAI program to confront an emerging vulnerability in modern artificial intelligence: the threat of AI T…
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
Anusha Sinha, Keltin Grimes, James Lucassen +4
A red team simulates adversary attacks to help defenders find effective strategies to defend their systems in a real-world operational setting. As more enterprise systems adopt AI,…
Concept-ROT: Poisoning Concepts in Large Language Models with Model Editing
Keltin Grimes, Marco Christiani, David Shriver +1
Model editing methods modify specific behaviors of Large Language Models by altering a small, targeted set of network weights and require very little data and compute. These method…