activity
20162022
most citedNon-Asymptotic Analysis of Robust Control from Coarse-Grained Identification

54 citations · 217 across the 12 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG20225 cited

On the Generalization of Representations in Reinforcement Learning

Charline Le Lan, Stephen Tu, Adam Oberman +2

In reinforcement learning, state representations are used to tractably deal with large problem spaces. State representations serve both to approximate the value function with few p…

cs.LG2020

Regret Bounds for Adaptive Nonlinear Control

Nicholas M. Boffi, Stephen Tu, Jean-Jacques E. Slotine

We study the problem of adaptively controlling a known discrete-time nonlinear system subject to unmodeled disturbances. We prove the first finite-time regret bounds for adaptive n…

cs.LG202028 cited

Learning Stability Certificates from Data

Nicholas M. Boffi, Stephen Tu, Nikolai Matni +2

Many existing tools in nonlinear control theory for establishing stability or safety of a dynamical system can be distilled to the construction of a certificate function that guara…

cs.LG201917 cited

Observational Overfitting in Reinforcement Learning

Xingyou Song, Yiding Jiang, Stephen Tu +2

A major component of overfitting in model-free reinforcement learning (RL) involves the case where the agent may mistakenly correlate reward with certain spurious features from the…

cs.LG201922 cited

Finite-time Analysis of Approximate Policy Iteration for the Linear Quadratic Regulator

Karl Krauth, Stephen Tu, Benjamin Recht

We study the sample complexity of approximate policy iteration (PI) for the Linear Quadratic Regulator (LQR), building on a recent line of work using LQR as a testbed to understand…

cs.LG2018

The Gap Between Model-Based and Model-Free Methods on the Linear Quadratic Regulator: An Asymptotic Viewpoint

Stephen Tu, Benjamin Recht

The effectiveness of model-based versus model-free methods is a long-standing question in reinforcement learning (RL). Motivated by recent empirical success of RL on continuous con…