391 citations · 478 across the 3 of their papers we have counts for
3 papers
cs.LG2022★ 87 cited
In-context Learning and Induction Heads
Catherine Olsson, Nelson Elhage, Neel Nanda +23
"Induction heads" are attention heads that implement a simple algorithm to complete token sequences like [A][B] ... [A] -> [B]. In this work, we present preliminary and indirect ev…
cs.CL2022★ 391 cited
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Yuntao Bai, Andy Jones, Kamal Ndousse +28
We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We find this alignment tra…
cs.LG2021
An Empirical Investigation of Learning from Biased Toxicity Labels
Neel Nanda, Jonathan Uesato, Sven Gowal
Collecting annotations from human raters often results in a trade-off between the quantity of labels one wishes to gather and the quality of these labels. As such, it is often only…