Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Scaling laws for activation steering with Llama 2 models and refusal mechanisms
Sheikh Abdur Raheem Ali, Justin Xu, Ivory Yang +3
As large language models (LLMs) evolve in complexity and capability, the efficacy of less widely deployed alignment techniques are uncertain. Building on previous work on activatio…
cs.LG2024
ProgressGym: Alignment with a Millennium of Moral Progress
Tianyi Qiu, Yang Zhang, Xuchuan Huang +3
Frontier AI systems, including large language models (LLMs), hold increasing influence over the epistemology of human users. Such influence can reinforce prevailing societal values…