5 papers
Mecha-nudges for Machines
Giulio Frey, Kawin Ethayarajh
AI agents are becoming active decision-makers on the Internet. As they make decisions in the same environments as humans, the environments themselves can change to influence them.…
Humanline: Online Alignment as Perceptual Loss
Sijia Liu, Niklas Muennighoff, Kawin Ethayarajh
Online alignment (e.g., GRPO) is generally more performant than offline alignment (e.g., DPO) -- but why? Drawing on prospect theory from behavioral economics, we propose a human-c…
Understanding Dataset Difficulty with -Usable Information
Kawin Ethayarajh, Yejin Choi, Swabha Swayamdipta
Estimating the difficulty of a dataset typically involves comparing state-of-the-art models to humans; the bigger the performance gap, the harder the dataset is said to be. However…
KTO: Model Alignment as Prospect Theoretic Optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff +2
Kahneman & Tversky's tells us that humans perceive random variables in a biased but well-defined manner (1992); for example, humans are famously loss-ave…
Data Checklist: On Unit-Testing Datasets with Usable Information
Heidi C. Zhang, Shabnam Behzad, Kawin Ethayarajh +1
Model checklists (Ribeiro et al., 2020) have emerged as a useful tool for understanding the behavior of LLMs, analogous to unit-testing in software engineering. However, despite da…