10 papers
A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training
Junze Ye, Jiayi Cheng, Miao Lu +3
For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels condition on. Pure behavioral cloning uses…
Annealed Softmax Greedy in Many-Armed Bayesian Bandits
William Overman, Mohsen Bayati
Reinforcement learning with verifiable rewards and group-based policy optimization methods update a stochastic policy by sampling multiple completions per prompt and increasing the…
Calibrating Conservatism for Scalable Oversight
William Overman, Mohsen Bayati
Agentic AI systems capable of autonomous planning and extended environmental interaction pose a fundamental control problem: how can humans maintain meaningful oversight of systems…
Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context
Yilun Zhu, Yuan Zhuang, Nikhita Vedula +6
Many applications of LLM-based text regression require predicting a full conditional distribution rather than a single point value. We study distributional regression under empiric…
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
William Overman, Mohsen Bayati
As increasingly capable agents are deployed, a central safety challenge is how to retain meaningful human control without modifying the underlying system. We study a minimal contro…
Scaling Clinician-Grade Feature Generation from Clinical Notes with Multi-Agent Language Models
Jiayi Wang, Jacqueline Jil Vallon, Nikhil V. Kotha +8
Developing accurate clinical prediction models is often bottlenecked by the difficulty of deriving meaningful structured features from unstructured EHR notes, a process that tradit…