3 papers
cs.AI2026
EigenBench: A Comparative Behavioral Measure of Value Alignment
Jonathn Chang, Leonhard Piff, Suvadip Sana +2
Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for compara…
cs.LG2025
Scaling laws for activation steering with Llama 2 models and refusal mechanisms
Sheikh Abdur Raheem Ali, Justin Xu, Ivory Yang +3
As large language models (LLMs) evolve in complexity and capability, the efficacy of less widely deployed alignment techniques are uncertain. Building on previous work on activatio…
cs.LG2024
ProgressGym: Alignment with a Millennium of Moral Progress
Tianyi Qiu, Yang Zhang, Xuchuan Huang +3
Frontier AI systems, including large language models (LLMs), hold increasing influence over the epistemology of human users. Such influence can reinforce prevailing societal values…