3 papers
cs.AI2025
Reliable Weak-to-Strong Monitoring of LLM Agents
Neil Kale, Chen Bo Calvin Zhang, Kevin Zhu +5
We stress test monitoring systems for detecting covert misbehavior in autonomous LLM agents (e.g., secretly sharing private information). To this end, we systematize a monitor red…
cs.CY2025
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
Christina Q. Knight, Kaustubh Deshpande, Ved Sirdeshmukh +4
The rapid advancement of large language models (LLMs) introduces dual-use capabilities that could both threaten and bolster national security and public safety (NSPS). Models imple…
cs.CL2025
Jailbreaking to Jailbreak
Jeremy Kritz, Vaughn Robinson, Robert Vacareanu +7
Large Language Models (LLMs) can be used to red team other models (e.g. jailbreaking) to elicit harmful contents. While prior works commonly employ open-weight models or private un…