From the 1 of 1 linked paper with an AI index.
1 paper
Huihao Jing, Haozhe Cui, Wenbin Hu +9
The paper introduces RLPF, a reinforcement‑learning approach that uses staged performance feedback to train code‑generation models to produce not only correct programs but also fas…