3 papers
cs.AI2026
PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost
Junkeun Yi, Damon Mosk-Aoyama, Baihe Huang +9
Post-training for long-horizon agentic tasks has a tension between compute efficiency and generalization. While supervised fine-tuning (SFT) is compute efficient, it often suffers…
cs.LG2026
Towards Anytime-Valid Statistical Watermarking
Baihe Huang, Eric Xu, Kannan Ramchandran +2
The proliferation of Large Language Models (LLMs) necessitates efficient mechanisms to distinguish machine-generated content from human text. While statistical watermarking has eme…
cs.LG2025
Sample Complexity and Representation Ability of Test-time Scaling Paradigms
Baihe Huang, Shanda Li, Tianhao Wu +5
Test-time scaling paradigms have significantly advanced the capabilities of large language models (LLMs) on complex tasks. Despite their empirical success, theoretical understandin…