16 papers
Backdoor Decontamination Dynamics in LLM Agents
Gabriel Huang, Abhay Puri, Léo Boisvert +4
Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing. Assuming defender…
PrivacyAlign: Contextual Privacy Alignment for LLM Agents
Manveer Singh Tamber, Abhay Puri, Marc-Etienne Brunet +3
AI agents acting on behalf of users are constantly making decisions, and for users to trust their agents, those decisions must align with what they actually want. Privacy is an imp…
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
Léo Boisvert, Léo Boisvert, Abhay Puri +8
While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces critical security vulnerabilities within the a…
Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks?
Rishika Bhagwatkar, Kevin Kasa, Abhay Puri +5
AI agents are vulnerable to indirect prompt injection attacks, where malicious instructions embedded in external content or tool outputs cause unintended or harmful behavior. Inspi…
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
Juan Rodriguez, Haotian Zhang, Abhay Puri +13
We introduce VectorGym, a comprehensive benchmark suite for Scalable Vector Graphics (SVG) that spans generation from text and sketches, complex editing, and visual understanding.…
Scale Dependent Data Duplication
Joshua Kazdan, Noam Levi, Rylan Schaeffer +6
Data duplication during pretraining can degrade generalization and lead to memorization, motivating aggressive deduplication pipelines. However, at web scale, it is unclear what co…