activity
20242026
collaborators

10 papers

cs.LG2026

A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training

Junze Ye, Jiayi Cheng, Miao Lu +3

For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels condition on. Pure behavioral cloning uses…

cs.LG2026

Annealed Softmax Greedy in Many-Armed Bayesian Bandits

William Overman, Mohsen Bayati

Reinforcement learning with verifiable rewards and group-based policy optimization methods update a stochastic policy by sampling multiple completions per prompt and increasing the…

cs.AI2026

Calibrating Conservatism for Scalable Oversight

William Overman, Mohsen Bayati

Agentic AI systems capable of autonomous planning and extended environmental interaction pose a fundamental control problem: how can humans maintain meaningful oversight of systems…

cs.CL2026

Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context

Yilun Zhu, Yuan Zhuang, Nikhita Vedula +6

Many applications of LLM-based text regression require predicting a full conditional distribution rather than a single point value. We study distributional regression under empiric…

cs.AI2026

The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy

William Overman, Mohsen Bayati

As increasingly capable agents are deployed, a central safety challenge is how to retain meaningful human control without modifying the underlying system. We study a minimal contro…

cs.AI2025

Scaling Clinician-Grade Feature Generation from Clinical Notes with Multi-Agent Language Models

Jiayi Wang, Jacqueline Jil Vallon, Nikhil V. Kotha +8

Developing accurate clinical prediction models is often bottlenecked by the difficulty of deriving meaningful structured features from unstructured EHR notes, a process that tradit…