activity
20192022
most citedTraining a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

391 citations · 513 across the 3 of their papers we have counts for

collaborators

5 papers

cs.HC202235 cited

Measuring Progress on Scalable Oversight for Large Language Models

Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez +43

Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on m…

cs.LG202287 cited

In-context Learning and Induction Heads

Catherine Olsson, Nelson Elhage, Neel Nanda +23

"Induction heads" are attention heads that implement a simple algorithm to complete token sequences like [A][B] ... [A] -> [B]. In this work, we present preliminary and indirect ev…

cs.CL2022391 cited

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Yuntao Bai, Andy Jones, Kamal Ndousse +28

We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We find this alignment tra…

gr-qc2020

Gravitational-wave physics with Cosmic Explorer: limits to low-frequency sensitivity

Evan D. Hall, Kevin Kuns, Joshua R. Smith +17

Cosmic Explorer (CE) is a next-generation ground-based gravitational-wave observatory concept, envisioned to begin operation in the 2030s, and expected to be capable of observing b…

quant-ph2019

A phase-sensitive optomechanical amplifier for quantum noise reduction in laser interferometers

Yuntao Bai, Gautam Venugopalan, Kevin Kuns +5

The sensitivity of future gravitational wave interferometers is expected to be limited through-out the detection band by quantum vacuum fluctuations, which can be reduced by quantu…