2 papers
cs.CR2026
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
Itay Zloczower, Eyal Lenga, Gilad Gressel +1
Model providers increasingly release open weights or allow users to fine-tune foundation models through APIs. Although these models are safety-aligned before release, their safegua…
cs.AI2026
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
Shir Rozenfeld, Rahul Pankajakshan, Itay Zloczower +3
Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be apparent at the surface-text level. Ho…