1 paper · 1 filter
Federico Barbero, Xiangming Gu, Christopher A. Choquette-Choo +6
In this work, we show that it is possible to extract significant amounts of alignment training data from a post-trained model -- useful to steer the model to improve certain capabi…