activity
20162025
most citedPlanning in Observable POMDPs in Quasipolynomial Time

4 citations · 6 across the 12 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2025

To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning

Yuda Song, Dhruv Rohatgi, Aarti Singh +1

Partial observability is a notorious challenge in reinforcement learning (RL), due to the need to learn complex, history-dependent policies. Recent empirical successes have used pr…

cs.LG2025

Taming Imperfect Process Verifiers: A Sampling Perspective on Backtracking

Dhruv Rohatgi, Abhishek Shetty, Donya Saless +4

Test-time algorithms that combine the generative power of language models with process verifiers that assess the quality of partial generations offer a promising lever for elicitin…

cs.LG2024

Online Control in Population Dynamics

Noah Golowich, Elad Hazan, Zhou Lu +2

The study of population dynamics originated with early sociological works but has since extended into many fields, including biology, epidemiology, evolutionary game theory, and ec…

cs.LG2024

Exploration is Harder than Prediction: Cryptographically Separating Reinforcement Learning from Supervised Learning

Noah Golowich, Ankur Moitra, Dhruv Rohatgi

Supervised learning is often computationally easy in practice. But to what extent does this mean that other modes of learning, such as reinforcement learning (RL), ought to be comp…

cs.LG2023

Exploring and Learning in Sparse Linear MDPs without Computationally Intractable Oracles

Noah Golowich, Ankur Moitra, Dhruv Rohatgi

The key assumption underlying linear Markov Decision Processes (MDPs) is that the learner has access to a known feature map that maps state-action pairs to -dimensiona…

cs.LG2023

Provable benefits of score matching

Chirag Pabbaraju, Dhruv Rohatgi, Anish Sevekari +3

Score matching is an alternative to maximum likelihood (ML) for estimating a probability distribution parametrized up to a constant of proportionality. By fitting the ''score'' of…