1 paper
James Lucassen, Mark Henry, Philippa Wright +1
Many theoretical obstacles to AI alignment are consequences of reflective stability - the problem of designing alignment mechanisms that the AI would not disable if given the optio…