3 papers
cs.LG2026
Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search
Ayushman Singh, Siddharth Aphale
Good action rankings do not make a contrastive critic safe to maximize. These critics increasingly act as value-like objectives for best-of- selection, planning, and critic-guid…
cs.LG2026
SCOUT: Per-Context Reset Curricula for Sparse-Reward Reinforcement Learning
Siddharth Aphale, Ayushman Singh
Sparse-reward reinforcement learning often fails because rollouts from the unassisted evaluation start rarely reach later task stages. Reset curricula address this by starting some…
cs.LG2026
SFT Overtraining Predicts Rank Inversion via Entropy Collapse Under RLVR
Siddharth Aphale, Kelly Liu
The standard heuristic of selecting the SFT checkpoint with the highest pass@1 for GRPO can fail when SFT compresses the rollout distribution. For binary rewards, the expected with…