Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
Nevan Wichers, Aram Ebtekar, Ariana Azarbal +8
Large language models are sometimes trained with imperfect oversight signals, leading to undesired behaviors such as reward hacking and sycophancy. Improving oversight quality can…
cs.LG2024
Visualizing Neural Network Imagination
Nevan Wichers, Victor Tao, Riccardo Volpato +1
In certain situations, neural networks will represent environment states in their hidden activations. Our goal is to visualize what environment states the networks are representing…