1 paper · 1 filter
Callum Canavan, Aditya Shrivastava, Allison Qi +2
To steer language models towards truthful outputs on tasks which are beyond human capability, previous work has suggested training models on easy tasks to steer them on harder ones…