activity
20212026
most citedConstitutional AI: Harmlessness from AI Feedback

334 citations · 510 across the 8 of their papers we have counts for

collaborators
Showing 2022Show all

5 papers · 1 filter

cs.CL2022★ 52 cited

Discovering Language Model Behaviors with Model-Written Evaluations

Ethan Perez, Sam Ringer, Kamilė Lukošiūtė +60

As language models (LMs) scale, they develop many novel behaviors, good and bad, exacerbating the need to evaluate how they behave. Prior work creates evaluations with crowdwork (w…

cs.CL2022★ 334 cited

Constitutional AI: Harmlessness from AI Feedback

Yuntao Bai, Saurav Kadavath, Sandipan Kundu +48

As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improveme…

cs.HC2022★ 35 cited

Measuring Progress on Scalable Oversight for Large Language Models

Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez +43

Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on m…

astro-ph.HE2022★ 9 cited

Prospects of Gravitational Wave Follow-up Through a Wide-field Ultra-violet Satellite: a Dorado Case Study

Bas Dorsman, Geert Raaijmakers, S. Bradley Cenko +8

The detection of gravitational waves from binary neuron star merger GW170817 and electromagnetic counterparts GRB170817 and AT2017gfo kick-started the field of gravitational wave m…

astro-ph.IM2022★ 16 cited

KilonovaNet: Surrogate Models of Kilonova Spectra with Conditional Variational Autoencoders

Kamilė Lukošiūtė, Geert Raaijmakers, Zoheyr Doctor +2

Detailed radiative transfer simulations of kilonova spectra play an essential role in multimessenger astrophysics. Using the simulation results in parameter inference studies requi…