3 papers
cs.LG2026
Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety
Domenic Rosati, Ali Dadsetan, Hong Huang +5
A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produce hate speech. Preventing su…
math.OC2026
Limits of Convergence-Rate Control for Open-Weight Safety
Domenic Rosati, Xijie Zeng, Hong Huang +4
Open-weight foundation models can be fine-tuned for harmful purposes after release, yet no existing training resistance methods provide theoretical guarantees. Treating these inter…
cs.LG2024
Evaluating Defences against Unsafe Feedback in RLHF
Domenic Rosati, Giles Edkins, Harsh Raj +5
While there has been progress towards aligning Large Language Models (LLMs) with human values and ensuring safe behaviour at inference time, safety guards can easily be removed whe…