3 papers
cs.CL2026
On Safety Risks in Experience-Driven Self-Evolving Agents
Weixiang Zhao, Yichen Zhang, Yingshuo Wang +8
Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self-curated experience introduc…
cs.LG2025
Validating Mechanistic Interpretations: An Axiomatic Approach
Nils Palumbo, Ravi Mangal, Zifan Wang +3
Mechanistic interpretability aims to reverse engineer the computation performed by a neural network in terms of its internal components. Although there is a growing body of researc…
cs.CR2024
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
Priyanshu Kumar, Elaine Lau, Saranya Vijayakumar +9
For safety reasons, large language models (LLMs) are trained to refuse harmful user instructions, such as assisting dangerous activities. We study an open question in this work: do…