2 papers
cs.LG2024
SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking
Chris Cundy, Stefano Ermon
In many domains, autoregressive models can attain high likelihood on the task of predicting the next observation. However, this maximum-likelihood (MLE) objective does not necessar…
cs.LG2024
Privacy-Constrained Policies via Mutual Information Regularized Policy Gradients
Chris Cundy, Rishi Desai, Stefano Ermon
As reinforcement learning techniques are increasingly applied to real-world decision problems, attention has turned to how these algorithms use potentially sensitive information. W…