3 papers
cs.LG2026
Overtrained, Not Misaligned
Joel Schreiber, Ariel Goldstein
Emergent misalignment (EM), where fine-tuning on a narrow task (like insecure code) causes broad misalignment across unrelated domains, was first demonstrated by Betley et al. (202…
cs.CL2026
Evaluating Alignment of Behavioral Dispositions in LLMs
Amir Taubenfeld, Zorik Gekhman, Lior Nezry +8
As LLMs integrate into our daily lives, understanding their behavior becomes essential. In this work, we focus on behavioral dispositionsthe underlying tendencies that shape res…
cs.HC2025
Confident-Knowledge Diversity Drives Human-Human and Human-AI Free Discussion Synergy and Reveals Pure-AI Discussion Shortfalls
Tom Sheffer, Alon Miron, Asael Sklar +2
Conversations transform individual knowledge into collective insight, enabling collaborators to solve problems more accurately than they could alone. Whether dialogues among large…