◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

A. Rakhlin

5 papers hereh-index 5812.5k citations176 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • last author5

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • math.PR1
  • math.ST1
  • stat.ML1

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.LGShow all

2 papers · 1 filter

cs.LG2024

Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Tengyang Xie, Dylan J. Foster, Akshay Krishnamurthy +3

Reinforcement learning from human feedback (RLHF) has emerged as a central tool for language model alignment. We consider online exploration in RLHF, which exploits interactive acc…

cs.LG2024

The Power of Resets in Online Reinforcement Learning

Zakaria Mhammedi, Dylan J. Foster, Alexander Rakhlin

Simulators are a pervasive tool in reinforcement learning, but most existing algorithms cannot efficiently exploit simulator access -- particularly in high-dimensional domains that…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.