1 paper
Manish Bhattarai, Ismael Boureima, Nishath Rajiv Ranasinghe +2
We argue that decomposing reward into weighted, verifiable criteria and using an LLM judge to score them provides a partial-credit optimization signal: instead of a binary outcome…