1 paper
Tianyi Alex Qiu, Micah Carroll, Cameron Allen
The evaluation and post-training of large language models (LLMs) rely on supervision, but strong supervision for difficult tasks is often unavailable, especially when evaluating fr…