3 papers
cs.CL2025
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
Taylor Sorensen, Benjamin Newman, Jared Moore +5
Language model post-training has enhanced instruction-following and performance on many downstream tasks, but also comes with an often-overlooked cost on tasks with many possible v…
cs.AI2024
Intuitions of Compromise: Utilitarianism vs. Contractualism
Jared Moore, Yejin Choi, Sydney Levine
What is the best compromise in a situation where different people value different things? The most commonly accepted method for answering this question -- in fields across the beha…
cs.CL2024
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
Yuling Gu, Oyvind Tafjord, Hyunwoo Kim +4
Large language models (LLMs) are increasingly tested for a "Theory of Mind" (ToM) - the ability to attribute mental states to oneself and others. Yet most evaluations stop at expli…