4 papers
Mixing Times of Glauber Dynamics on Masked Language Models
Suvadip Sana, Sami Wolf, Neer Mehta +4
Masked language models (MLMs) define local conditional distributions over tokens but do not, in general, correspond to any consistent joint distribution over sequences. This raises…
EigenBench: A Comparative Behavioral Measure of Value Alignment
Jonathn Chang, Leonhard Piff, Suvadip Sana +2
Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for compara…
Democratic Preference Alignment via Sortition-Weighted RLHF
Suvadip Sana, Jinzhou Wu, Martin T. Wells
Whose values should AI systems learn? Preference based alignment methods like RLHF derive their training signal from human raters, yet these rater pools are typically convenience s…
Quantitative Relaxations of Arrow's Axioms
Suvadip Sana, Daniel Brous, Martin T. Wells +1
In this paper we develop a novel approach to relaxing Arrow's axioms for voting rules, addressing a long-standing critique in social choice theory. Classical axioms (often styled a…