From the 1 of 1 linked paper with an AI index.
1 paper
Tiangang Li, Xiangbo Tian
The paper introduces HARGO, a reinforcement‑learning post‑training method that weights responses by confidence and reward contrast to better align large language models with divers…