3 papers
cs.CR2026
ContainmentBench: Trace-Based Evaluation of Post-Exposure Containment in Tool-Using LLM Agents
Wenhao Lan, Shan Li, Xinhua Lai +3
Tool-using large language model (LLM) agents read untrusted content, maintain memory, delegate tasks, and invoke tools with external side effects. Terminal attack-success or policy…
cs.CR2026
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning
Wenhao Lan, Shan Li, Xinhua Lai +3
Safety alignment requires language models to refuse harmful requests without losing the ability to answer benign ones. Existing robustness evaluations, however, do not reveal wheth…
cs.LG2026
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
Wenhao Lan, Shan Li, Xinhua Lai +4
Safety-aligned language models must refuse harmful requests without broad over-refusal, but it remains unclear how dynamic adversarial fine-tuning changes refusal-control carriers:…