action ranking 1adaptive scheduling 1bilinear critics 1contrastive learning 1curriculum learning 1norm drift 1policy search 1reinforcement learning 1reset curricula 1sparse rewards 1value calibration 1
From the 2 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search
Ayushman Singh, Siddharth Aphale
The paper investigates why bilinear contrastive critics, which rank actions for reinforcement learning policies, can produce unsafe or misleading rankings due to issues like norm d…
cs.LG2026
SCOUT: Per-Context Reset Curricula for Sparse-Reward Reinforcement Learning
Siddharth Aphale, Ayushman Singh
SCOUT is a reset curriculum method for sparse-reward reinforcement learning that adaptively provides per-context assistance and removes it based on binary rollout success, without…