1 citations · 1 across the 2 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
Nevan Wichers, Aram Ebtekar, Ariana Azarbal +8
Large language models are sometimes trained with imperfect oversight signals, leading to undesired behaviors such as reward hacking and sycophancy. Improving oversight quality can…
cs.LG2020
Resolving Spurious Correlations in Causal Models of Environments via Interventions
Sergei Volodin, Nevan Wichers, Jeremy Nixon
Causal models bring many benefits to decision-making systems (or agents) by making them interpretable, sample-efficient, and robust to changes in the input distribution. However, s…