#reward design

topicreward design

4 papers · 1 filter

cs.LG2026

Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning

Efstratios Zaradoukas, Davide Gabrielli, Bardh Prenkaj +1

The paper investigates how different reward functions affect the speed and effectiveness of reinforcement‑learning based machine unlearning for language models, proposing graded an…

cs.LG2026

Reinforcement Learning for Code Optimization

Pierre Chambon, Kunhao Zheng, Juliette Decugis +2

The paper proposes a reinforcement‑learning framework that learns to optimize program execution speed by addressing measurement noise, sparse rewards, and instability, using a cali…

cs.CV2026

Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Rui Li, Yuanzhi Liang, Ziqi Ni +3

The paper proposes TaRoS, a framework that redesigns reward signals for video generation using GRPO to avoid reward hacking and saturation, improving visual fidelity, motion cohere…

cs.LG2026

When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR

Chuyifei Zhang

The paper investigates how test suites used as rewards for reinforcement‑learning‑based code generation can contain systematic false positives, and shows that hardening the suite r…