16 citations · 30 across the 4 of their papers we have counts for
4 papers
Training Language Models with Language Feedback at Scale
Jérémy Scheurer, Jon Ander Campos, Tomasz Korbak +4
Pretrained language models often generate outputs that are not in line with human preferences, such as harmful text or factually incorrect summaries. Recent work approaches the abo…
How Would The Viewer Feel? Estimating Wellbeing From Video Scenarios
Mantas Mazeika, Eric Tang, Andy Zou +6
In recent years, deep neural networks have demonstrated increasingly strong abilities to recognize objects and activities in videos. However, as video understanding becomes widely…
Few-shot Adaptation Works with UnpredicTable Data
Jun Shern Chan, Michael Pieler, Jonathan Jao +2
Prior work on language models (LMs) shows that training on a large number of diverse tasks improves few-shot learning (FSL) performance on new tasks. We take this to the extreme, a…
Training Language Models with Language Feedback
Jérémy Scheurer, Jon Ander Campos, Jun Shern Chan +3
Pretrained language models often do not perform tasks in ways that are in line with our preferences, e.g., generating offensive text or factually incorrect summaries. Recent work a…