activity
20232026
most citedGuardrail Baselines for Unlearning in LLMs

4 citations · 8 across the 7 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL

Zhoujun Cheng, Yutao Xie, Yuxiao Qu +12

While scaling laws guide compute allocation for LLM pre-training, analogous prescriptions for reinforcement learning (RL) post-training of large language models (LLMs) remain poorl…

cs.LG2026

POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration

Yuxiao Qu, Amrith Setlur, Virginia Smith +2

Reinforcement learning (RL) has improved the reasoning abilities of large language models (LLMs), yet state-of-the-art methods still fail to learn on many training problems. On har…

cs.LG2025

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Amrith Setlur, Matthew Y. R. Yang, Charlie Snell +5

Test-time scaling offers a promising path to improve LLM reasoning by utilizing more compute at inference time; however, the true promise of this paradigm lies in extrapolation (i.…

cs.LG2024★ 1 cited

RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Amrith Setlur, Saurabh Garg, Xinyang Geng +3

Training on model-generated synthetic data is a promising approach for finetuning LLMs, but it remains unclear when it helps or hurts. In this paper, we investigate this question f…

cs.LG2023★ 3 cited

Complementary Benefits of Contrastive Learning and Self-Training Under Distribution Shift

Saurabh Garg, Amrith Setlur, Zachary Chase Lipton +3

Self-training and contrastive learning have emerged as leading techniques for incorporating unlabeled data, both under distribution shift (unsupervised domain adaptation) and when…

cs.LG2023

On the Benefits of Public Representations for Private Transfer Learning under Distribution Shift

Pratiksha Thaker, Amrith Setlur, Zhiwei Steven Wu +1

Public pretraining is a promising approach to improve differentially private model training. However, recent work has noted that many positive research results studying this paradi…