8 citations · 10 across the 2 of their papers we have counts for
2 papers
cs.CL2024★ 2 cited
Prover-Verifier Games improve legibility of LLM outputs
Jan Hendrik Kirchner, Yining Chen, Harri Edwards +3
One way to increase confidence in the outputs of Large Language Models (LLMs) is to support them with reasoning that is clear and easy to check -- a property we call legibility. We…
cs.SE2024★ 8 cited
LLM Critics Help Catch LLM Bugs
Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe +3
Reinforcement learning from human feedback (RLHF) is fundamentally limited by the capacity of humans to correctly evaluate model output. To improve human evaluation ability and ove…