1 paper · 1 filter
Eric Wallace, Olivia Watkins, Miles Wang +2
In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning…