1 paper
Ethan Leung, Elias Lumer, Corey Feld +3
Reinforcement learning increasingly relies on an LLM judge to score each rubric criterion, and that judge acts as the reward model during training. Before such a signal can be trus…