◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Neel Nanda

3 papers hereh-index 3711.6k citations74 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author2

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • cs.CL1

identity via Semantic Scholar / OpenAlex

most citedTraining a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

391 citations · 478 across the 3 of their papers we have counts for

collaborators

3 papers

cs.LG2022★ 87 cited

In-context Learning and Induction Heads

Catherine Olsson, Nelson Elhage, Neel Nanda +23

"Induction heads" are attention heads that implement a simple algorithm to complete token sequences like [A][B] ... [A] -> [B]. In this work, we present preliminary and indirect ev…

cs.CL2022★ 391 cited

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Yuntao Bai, Andy Jones, Kamal Ndousse +28

We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We find this alignment tra…

cs.LG2021

An Empirical Investigation of Learning from Biased Toxicity Labels

Neel Nanda, Jonathan Uesato, Sven Gowal

Collecting annotations from human raters often results in a trade-off between the quantity of labels one wishes to gather and the quality of these labels. As such, it is often only…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.