1 citations · 1 across the 4 of their papers we have counts for
5 papers · 1 filter
An Empirical Risk Minimization Approach for Offline Inverse RL and Dynamic Discrete Choice Model
Enoch H. Kang, Hema Yoganarasimhan, Lalit Jain
We study the problem of estimating Dynamic Discrete Choice (DDC) models, also known as offline Maximum Entropy-Regularized Inverse Reinforcement Learning (offline MaxEnt-IRL) in ma…
A Lecture Note on Offline RL and IRL, Part II: Foundations of Inverse Reinforcement Learning and Dynamic Discrete Choice Models
Enoch Hyunwook Kang
In the forward reinforcement-learning problem, the reward is fixed and known; the learner is asked to find a good policy or value function. Here we turn the question around. Given…
Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity
Enoch Hyunwook Kang
Personalized alignment aims to adapt large language models to heterogeneous user preferences, yet the precise theoretical conditions for its statistical efficiency have not been fo…
Demystifying the unreasonable effectiveness of online alignment methods
Enoch Hyunwook Kang
Iterative alignment methods based on purely greedy updates are remarkably effective in practice, yet existing theoretical guarantees of \(O(\log T)\) KL-regularized regret can seem…
Stability and Generalization for Bellman Residuals
Enoch H. Kang, Kyoungseok Jang
Offline reinforcement learning and offline inverse reinforcement learning aim to recover near-optimal value functions or reward models from a fixed batch of logged trajectories, ye…