Showing cs.CRShow all
2 papers · 1 filter
cs.CR2026
ContainmentBench: Trace-Based Evaluation of Post-Exposure Containment in Tool-Using LLM Agents
Wenhao Lan, Shan Li, Xinhua Lai +3
Tool-using large language model (LLM) agents read untrusted content, maintain memory, delegate tasks, and invoke tools with external side effects. Terminal attack-success or policy…
cs.CR2026
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning
Wenhao Lan, Shan Li, Xinhua Lai +3
Safety alignment requires language models to refuse harmful requests without losing the ability to answer benign ones. Existing robustness evaluations, however, do not reveal wheth…