1 paper · 1 filter
Chonghe Jiang, Bao Nguyen, Anthony Man-Cho So +1
Language models (LMs) can produce texts that appear accurate and coherent but contain untruthful or toxic content. Inference-time interventions that edit the hidden activations hav…