4 papers
Argument Collapse: LLMs Flatten Long-Form Public Debate
Yekyung Kim, Yapei Chang, Chau Minh Pham +1
As LLMs are increasingly used to draft public-facing arguments, they may flatten public debate by repeatedly introducing the same polished, plausible arguments. We study argument c…
Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs
Vinay Samuel, Yapei Chang, Mohit Iyyer
Many open-ended instructions have multiple valid answers that users can benefit from seeing, but post-training often narrows an LLM's output space toward a small set of canonical r…
Auto-Eval Judge: Towards a General Agentic Framework for Task Completion Evaluation
Roshita Bhonsle, Rishav Dutta, Sneha Vavilapalli +8
The increasing adoption of foundation models as agents across diverse domains necessitates a robust evaluation framework. Current methods, such as LLM-as-a-Judge, focus only on fin…
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
Zongxia Li, Yapei Chang, Yuhang Zhou +4
Evaluating open-ended long-form generation is challenging because it is hard to define what clearly separates good from bad outputs. Existing methods often miss key aspects like co…