advantage shaping 1contrastive learning 1on-policy distillation 1policy optimization 1token-level correctness 1
From the 1 of 4 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
Weiwen Xu, Jia Liu, Hou Pong Chan +4
The paper proposes Contrastive Policy Optimization, which leverages token‑level contrastive disagreement between reference‑guided and standard generation distributions to provide a…
cs.LG2025
Resource Efficient Sleep Staging via Multi-Level Masking and Prompt Learning
Lejun Ai, Yulong Li, Haodong Yi +5
Automatic sleep staging plays a vital role in assessing sleep quality and diagnosing sleep disorders. Most existing methods rely heavily on long and continuous EEG recordings, whic…