13 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.LG2025
Trading Inference-Time Compute for Adversarial Robustness
Wojciech Zaremba, Evgenia Nitishinskaya, Boaz Barak +8
We conduct experiments on the impact of increasing inference-time compute in reasoning models (specifically OpenAI o1-preview and o1-mini) on their robustness to adversarial attack…
cs.CL2025★ 13 cited
Deliberative Alignment: Reasoning Enables Safer Language Models
Melody Y. Guan, Manas Joglekar, Eric Wallace +12
As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge. We introdu…