From the 1 of 1 linked paper with an AI index.
1 paper
Shuhang Wang, Ziming Li, Hui Cheng
The paper introduces DHRCL, a reinforcement‑learning framework for code‑focused large language models that uses a hierarchy of dense rewards (syntax, execution, unit‑test pass, and…