2 papers
cs.AI2026
BASIL: Bayesian Assessment of Sycophancy in LLMs
Katherine Atwell, Pedram Heydari, Anthony Sicilia +1
Sycophancy (overly agreeable or flattering behavior) poses a fundamental challenge for human-AI collaboration, particularly in high-stakes decision-making domains such as health, l…
cs.LG2026
MixDPO: Modeling Preference Strength for Pluralistic Alignment
Saki Imai, Pedram Heydari, Anthony Sicilia +3
Preference based alignment objectives implicitly assume that all human preferences are expressed with equal strength. In practice, however, preference strength varies across indivi…