1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Dan Zhang, Min Cai, Jonathan Light +3
Reward models are central to both reinforcement learning (RL) with language models and inference-time verification. However, existing reward models often lack temporal consistency,…