17 citations · 17 across the 1 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Estimating Worst-Case Frontier Risks of Open-Weight LLMs
Eric Wallace, Olivia Watkins, Miles Wang +2
In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning…
cs.LG2025
Trading Inference-Time Compute for Adversarial Robustness
Wojciech Zaremba, Evgenia Nitishinskaya, Boaz Barak +8
We conduct experiments on the impact of increasing inference-time compute in reasoning models (specifically OpenAI o1-preview and o1-mini) on their robustness to adversarial attack…