◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Saurav Kadavath

UC Berkeley

20 papers hereh-index 1621.6k citations42 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3
  • middle author11

Across the 14 of 20 papers where every author was matched, so the position is known.

fields
  • cs.CL10
  • cs.LG7
  • cs.AI1
  • cs.CV1
  • cs.SE1
affiliations
  • UC Berkeley
Homepage

identity via Semantic Scholar / OpenAlex

activity
20192026
most citedTraining a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

391 citations · 1.3k across the 17 of their papers we have counts for

collaborators
Showing 2025 · cs.CLShow all

2 papers · 2 filters

cs.CL2025

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning

Congmin Zheng, Jiachen Zhu, Jianghao Lin +6

Process Reward Models (PRMs) play a central role in evaluating and guiding multi-step reasoning in large language models (LLMs), especially for mathematical problem solving. Howeve…

cs.CL2025

Jailbreak Distillation: Renewable Safety Benchmarking

Jingyu Zhang, Ahmed Elgohary, Xiawei Wang +5

Large language models (LLMs) are rapidly deployed in critical applications, raising urgent needs for robust safety benchmarking. We propose Jailbreak Distillation (JBDistill), a no…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.