8 citations · 9 across the 3 of their papers we have counts for
3 papers
cs.CL2025
On the Robustness of Reward Models for Language Model Alignment
Jiwoo Hong, Noah Lee, Eunki Kim +5
The Bradley-Terry (BT) model is widely practiced in reward modeling for reinforcement learning with human feedback (RLHF). Despite its effectiveness, reward models (RMs) trained wi…
cs.CL2024★ 1 cited
Evaluating the Consistency of LLM Evaluators
Noah Lee, Jiwoo Hong, James Thorne
Large language models (LLMs) have shown potential as general evaluators along with the evident benefits of speed and cost. While their correlation against human annotators has been…
cs.CL2024★ 8 cited
ORPO: Monolithic Preference Optimization without Reference Model
Jiwoo Hong, Noah Lee, James Thorne
While recent preference alignment algorithms for language models have demonstrated promising results, supervised fine-tuning (SFT) remains imperative for achieving successful conve…