◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Robert Dadashi

8 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author5
  • last author1

Across the 6 of 8 papers where every author was matched, so the position is known.

fields
  • cs.LG4
  • cs.CL3
  • cs.RO1
same name
  • Robert Dadashi — 11 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedGemma: Open Models Based on Gemini Research and Technology

238 citations · 392 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2024★ 1 cited

BOND: Aligning LLMs with Best-of-N Distillation

Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot +17

Reinforcement learning from human feedback (RLHF) is a key driver of quality and safety in state-of-the-art large language models. Yet, a surprisingly simple and strong inference-t…

cs.LG2024★ 1 cited

WARP: On the Benefits of Weight Averaged Rewarded Policies

Alexandre Ramé, Johan Ferret, Nino Vieillard +7

Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) by encouraging their generations to have high rewards, using a reward model trained on human p…

cs.LG2024★ 3 cited

WARM: On the Benefits of Weight Averaged Reward Models

Alexandre Ramé, Nino Vieillard, Léonard Hussenot +4

Aligning large language models (LLMs) with human preferences through reinforcement learning (RLHF) can lead to reward hacking, where LLMs exploit failures in the reward model (RM)…

cs.LG2023

Offline Reinforcement Learning with On-Policy Q-Function Regularization

Laixi Shi, Robert Dadashi, Yuejie Chi +2

The core challenge of offline reinforcement learning (RL) is dealing with the (potentially catastrophic) extrapolation error induced by the distribution shift between the history d…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.