activity
20162023
most citedEvaluating Large Language Models Trained on Code

1.5k citations · 1.9k across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL202353 cited

The Capacity for Moral Self-Correction in Large Language Models

Deep Ganguli, Amanda Askell, Nicholas Schiefer +46

We test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to "morally self-correct" -- to avoid producing harmf…

cs.CL2022119 cited

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Deep Ganguli, Liane Lovitt, Jackson Kernion +33

We describe our early efforts to red team language models in order to simultaneously discover, measure, and attempt to reduce their potentially harmful outputs. We make three main…

cs.CL2022170 cited

Language Models (Mostly) Know What They Know

Saurav Kadavath, Tom Conerly, Amanda Askell +33

We study whether language models can evaluate the validity of their own claims and predict which questions they will be able to answer correctly. We first show that larger models a…

cs.CL202127 cited

A General Language Assistant as a Laboratory for Alignment

Amanda Askell, Yuntao Bai, Anna Chen +19

Given the broad capabilities of large language models, it should be possible to work towards a general-purpose, text-based assistant that is aligned with human values, meaning that…

cs.CL201636 cited

Learning a Natural Language Interface with Neural Programmer

Arvind Neelakantan, Quoc V. Le, Martin Abadi +2

Learning a natural language interface for database tables is a challenging task that involves deep language understanding and multi-step reasoning. The task is often approached by…