3 citations · 5 across the 4 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024★ 2 cited
Prover-Verifier Games improve legibility of LLM outputs
Jan Hendrik Kirchner, Yining Chen, Harri Edwards +3
One way to increase confidence in the outputs of Large Language Models (LLMs) is to support them with reasoning that is clear and easy to check -- a property we call legibility. We…
cs.CL2023
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner +9
Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the ability of humans to supervise model behavior - for example, to evaluate wh…