13 citations · 34 across the 7 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025★ 1 cited
Training LLMs for Honesty via Confessions
Manas Joglekar, Jeremy Chen, Gabriel Wu +4
Large language models (LLMs) can be dishonest when reporting on their actions and beliefs -- for example, they may overstate their confidence in factual claims or cover up evidence…
cs.LG2025
Trading Inference-Time Compute for Adversarial Robustness
Wojciech Zaremba, Evgenia Nitishinskaya, Boaz Barak +8
We conduct experiments on the impact of increasing inference-time compute in reasoning models (specifically OpenAI o1-preview and o1-mini) on their robustness to adversarial attack…