4 citations · 6 across the 12 of their papers we have counts for
7 papers · 1 filter
To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning
Yuda Song, Dhruv Rohatgi, Aarti Singh +1
Partial observability is a notorious challenge in reinforcement learning (RL), due to the need to learn complex, history-dependent policies. Recent empirical successes have used pr…
Taming Imperfect Process Verifiers: A Sampling Perspective on Backtracking
Dhruv Rohatgi, Abhishek Shetty, Donya Saless +4
Test-time algorithms that combine the generative power of language models with process verifiers that assess the quality of partial generations offer a promising lever for elicitin…
Online Control in Population Dynamics
Noah Golowich, Elad Hazan, Zhou Lu +2
The study of population dynamics originated with early sociological works but has since extended into many fields, including biology, epidemiology, evolutionary game theory, and ec…
Exploration is Harder than Prediction: Cryptographically Separating Reinforcement Learning from Supervised Learning
Noah Golowich, Ankur Moitra, Dhruv Rohatgi
Supervised learning is often computationally easy in practice. But to what extent does this mean that other modes of learning, such as reinforcement learning (RL), ought to be comp…
Exploring and Learning in Sparse Linear MDPs without Computationally Intractable Oracles
Noah Golowich, Ankur Moitra, Dhruv Rohatgi
The key assumption underlying linear Markov Decision Processes (MDPs) is that the learner has access to a known feature map that maps state-action pairs to -dimensiona…
Provable benefits of score matching
Chirag Pabbaraju, Dhruv Rohatgi, Anish Sevekari +3
Score matching is an alternative to maximum likelihood (ML) for estimating a probability distribution parametrized up to a constant of proportionality. By fitting the ''score'' of…