2 papers
cs.AI2024
Cluster-norm for Unsupervised Probing of Knowledge
Walter Laurito, Sharan Maiya, Grégoire Dhimoïla +3
The deployment of language models brings challenges in generating reliable information, especially when these models are fine-tuned using human preferences. To extract encoded know…
cs.AI2024
Evaluating Stability of Unreflective Alignment
James Lucassen, Mark Henry, Philippa Wright +1
Many theoretical obstacles to AI alignment are consequences of reflective stability - the problem of designing alignment mechanisms that the AI would not disable if given the optio…