3 papers
cs.LG2026
TREK: Distill to Explore, Reinforce to Refine
Yuanda Xu, Zhengze Zhou, Kayhan Behdin +10
Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning trajectories, but it stalls on hard prompts whose correct solution m…
cs.DC2026
Large-Scale Regularized Matching on GPU Clusters
Aida Rahmattalabi, Gregory Dexter, Sanjana Garg +5
Production decision systems such as ad allocation or content matching involve millions of users and thousands of items, reducing to large-scale linear programs with sparse block-di…
cs.DC2026
DuaLip-GPU Technical Report
Gregory Dexter, Aida Rahmattalabi, Sanjana Garg +6
Large-scale linear programs (LPs) arise in many decision systems, including ranking, allocation, and matching problems that must be solved repeatedly at massive scale. Prior work s…