3 papers
cs.CR2025
DoomArena: A framework for Testing AI Agents Against Evolving Security Threats
Leo Boisvert, Mihir Bansal, Chandra Kiran Reddy Evuru +9
We present DoomArena, a security evaluation framework for AI agents. DoomArena is designed on three principles: 1) It is a plug-in framework and integrates easily into realistic ag…
cs.CR2025
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
Joshua Kazdan, Abhay Puri, Rylan Schaeffer +5
Leading language model (LM) providers like OpenAI and Anthropic allow customers to fine-tune frontier LMs for specific use cases. To prevent abuse, these providers apply filters to…
cs.CL2025
LitLLMs, LLMs for Literature Review: Are we there yet?
Shubham Agarwal, Gaurav Sahu, Abhay Puri +5
Literature reviews are an essential component of scientific research, but they remain time-intensive and challenging to write, especially due to the recent influx of research paper…