334 citations · 510 across the 8 of their papers we have counts for
5 papers · 1 filter
Discovering Language Model Behaviors with Model-Written Evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė +60
As language models (LMs) scale, they develop many novel behaviors, good and bad, exacerbating the need to evaluate how they behave. Prior work creates evaluations with crowdwork (w…
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu +48
As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improveme…
Measuring Progress on Scalable Oversight for Large Language Models
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez +43
Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on m…
Prospects of Gravitational Wave Follow-up Through a Wide-field Ultra-violet Satellite: a Dorado Case Study
Bas Dorsman, Geert Raaijmakers, S. Bradley Cenko +8
The detection of gravitational waves from binary neuron star merger GW170817 and electromagnetic counterparts GRB170817 and AT2017gfo kick-started the field of gravitational wave m…
KilonovaNet: Surrogate Models of Kilonova Spectra with Conditional Variational Autoencoders
Kamilė Lukošiūtė, Geert Raaijmakers, Zoheyr Doctor +2
Detailed radiative transfer simulations of kilonova spectra play an essential role in multimessenger astrophysics. Using the simulation results in parameter inference studies requi…