activity
20162026
most citedConditional Hardness of Earth Mover Distance

13 citations · 29 across the 26 of their papers we have counts for

collaborators
Showing 2025 · cs.LGShow all

5 papers · 2 filters

cs.LG2025

To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable Reinforcement Learning

Yuda Song, Dhruv Rohatgi, Aarti Singh +1

Partial observability is a notorious challenge in reinforcement learning (RL), due to the need to learn complex, history-dependent policies. Recent empirical successes have used pr…

cs.LG2025

Taming Imperfect Process Verifiers: A Sampling Perspective on Backtracking

Dhruv Rohatgi, Abhishek Shetty, Donya Saless +4

Test-time algorithms that combine the generative power of language models with process verifiers that assess the quality of partial generations offer a promising lever for elicitin…

cs.LG2025

Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration

Dylan J. Foster, Zakaria Mhammedi, Dhruv Rohatgi

Language model alignment (or, reinforcement learning) techniques that leverage active exploration -- deliberately encouraging the model to produce diverse, informative responses --…

cs.LG2025

Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification

Dhruv Rohatgi, Adam Block, Audrey Huang +2

Next-token prediction with the logarithmic loss is a cornerstone of autoregressive sequence modeling, but, in practice, suffers from error amplification, where errors in the model…

cs.LG2025

Necessary and Sufficient Oracles: Toward a Computational Taxonomy For Reinforcement Learning

Dhruv Rohatgi, Dylan J. Foster

Algorithms for reinforcement learning (RL) in large state spaces crucially rely on supervised learning subroutines to estimate objects such as value functions or transition probabi…