◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

J. Bagnell

4 papers hereh-index 6529.3k citations219 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2
  • last author2

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG4

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators

4 papers

cs.LG2026

Expanding the Capabilities of Reinforcement Learning via Text Feedback

Yuda Song, Lili Chen, Fahim Tajwar +5

The success of RL for LLM post-training stems from an unreasonably uninformative source: a single bit of information per rollout as binary reward or preference label. At the other…

cs.LG2025

All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning

Gokul Swamy, Sanjiban Choudhury, Wen Sun +2

From a first-principles perspective, it may seem odd that the strongest results in foundation model fine-tuning (FT) are achieved via a relatively complex, two-stage training proce…

cs.LG2025

To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning

Yuda Song, Dhruv Rohatgi, Aarti Singh +1

Partial observability is a notorious challenge in reinforcement learning (RL), due to the need to learn complex, history-dependent policies. Recent empirical successes have used pr…

cs.LG2024

REBEL: Reinforcement Learning via Regressing Relative Rewards

Zhaolin Gao, Jonathan D. Chang, Wenhao Zhan +7

While originally developed for continuous control problems, Proximal Policy Optimization (PPO) has emerged as the work-horse of a variety of reinforcement learning (RL) application…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.