◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

S. Kravec

3 papers hereh-index 1614k citations28 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL1
  • cs.HC1
  • cs.LG1

identity via Semantic Scholar / OpenAlex

most citedTraining a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

391 citations · 474 across the 3 of their papers we have counts for

collaborators

3 papers

cs.HC2022★ 35 cited

Measuring Progress on Scalable Oversight for Large Language Models

Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez +43

Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on m…

cs.LG2022★ 48 cited

Toy Models of Superposition

Nelson Elhage, Tristan Hume, Catherine Olsson +13

Neural networks often pack many unrelated concepts into a single neuron - a puzzling phenomenon known as 'polysemanticity' which makes interpretability much more challenging. This…

cs.CL2022★ 391 cited

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Yuntao Bai, Andy Jones, Kamal Ndousse +28

We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We find this alignment tra…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.