5 papers
Response Time Enhances Alignment with Heterogeneous Preferences
Federico Echenique, Alireza Fallah, Baihe Huang +1
Aligning large language models (LLMs) to human preferences typically relies on aggregating pooled feedback into a single reward model. However, this standard approach assumes that…
PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost
Junkeun Yi, Damon Mosk-Aoyama, Baihe Huang +9
Post-training for long-horizon agentic tasks has a tension between compute efficiency and generalization. While supervised fine-tuning (SFT) is compute efficient, it often suffers…
Towards Anytime-Valid Statistical Watermarking
Baihe Huang, Eric Xu, Kannan Ramchandran +2
The proliferation of Large Language Models (LLMs) necessitates efficient mechanisms to distinguish machine-generated content from human text. While statistical watermarking has eme…
Sample Complexity and Representation Ability of Test-time Scaling Paradigms
Baihe Huang, Shanda Li, Tianhao Wu +5
Test-time scaling paradigms have significantly advanced the capabilities of large language models (LLMs) on complex tasks. Despite their empirical success, theoretical understandin…
Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics
Hanlin Zhu, Baihe Huang, Shaolun Zhang +4
Auto-regressive large language models (LLMs) show impressive capacities to solve many complex reasoning tasks while struggling with some simple logical reasoning tasks such as inve…