4 papers
Better LLM Reasoning via Dual-Play
Zhengxin Zhang, Chengyu Huang, Aochong Oliver Li +1
Large Language Models (LLMs) have achieved remarkable progress through Reinforcement Learning with Verifiable Rewards (RLVR), yet still rely heavily on external supervision (e.g.,…
Judging with Confidence: Calibrating Autoraters to Preference Distributions
Zhuohang Li, Xiaowei Li, Chengyu Huang +11
The alignment of large language models (LLMs) with human values increasingly relies on using other LLMs as automated judges, or ``autoraters''. However, their reliability is limite…
DCRM: A Heuristic to Measure Response Pair Quality in Preference Optimization
Chengyu Huang, Tanya Goyal
Recent research has attempted to associate preference optimization (PO) performance with the underlying preference datasets. In this work, our observation is that the differences b…
HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
Chengyu Huang, Zhengxin Zhang, Claire Cardie
While scaling the length of responses at test-time has been shown to markedly improve the reasoning abilities and performance of large language models (LLMs), it often results in v…