activity
20182025
most citedWhy Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability

10 citations · 21 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

15 papers · 1 filter

cs.LG2025

Visual Pre-Training on Unlabeled Images using Reinforcement Learning

Dibya Ghosh, Sergey Levine

In reinforcement learning (RL), value-based algorithms learn to associate each observation with the states and rewards that are likely to be reached from it. We observe that many s…

cs.LG20252 cited

: a Vision-Language-Action Model with Open-World Generalization

Physical Intelligence, Kevin Black, Noah Brown +33

In order for robots to be useful, they must perform practically relevant tasks in the real world, outside of the lab. While vision-language-action (VLA) models have demonstrated im…

cs.LG2025

ViVa: Video-Trained Value Functions for Guiding Online RL from Diverse Data

Nitish Dashora, Dibya Ghosh, Sergey Levine

Online reinforcement learning (RL) with sparse rewards poses a challenge partly because of the lack of feedback on states leading to the goal. Furthermore, expert offline data with…

cs.LG2024

What Do Learning Dynamics Reveal About Generalization in LLM Reasoning?

Katie Kang, Amrith Setlur, Dibya Ghosh +4

Despite the remarkable capabilities of modern large language models (LLMs), the mechanisms behind their problem-solving abilities remain elusive. In this work, we aim to better und…

cs.LG2023

Accelerating Exploration with Unlabeled Prior Data

Qiyang Li, Jason Zhang, Dibya Ghosh +2

Learning to solve tasks from a sparse reward signal is a major challenge for standard reinforcement learning (RL) algorithms. However, in the real world, agents rarely need to solv…

cs.LG2023

HIQL: Offline Goal-Conditioned RL with Latent States as Actions

Seohong Park, Dibya Ghosh, Benjamin Eysenbach +1

Unsupervised pre-training has recently become the bedrock for computer vision and natural language processing. In reinforcement learning (RL), goal-conditioned RL can potentially p…