1 paper
Sheikh Abdur Raheem Ali, Justin Xu, Ivory Yang +3
As large language models (LLMs) evolve in complexity and capability, the efficacy of less widely deployed alignment techniques are uncertain. Building on previous work on activatio…