8 papers
An Empirical Risk Minimization Approach for Offline Inverse RL and Dynamic Discrete Choice Model
Enoch H. Kang, Hema Yoganarasimhan, Lalit Jain
We study the problem of estimating Dynamic Discrete Choice (DDC) models, also known as offline Maximum Entropy-Regularized Inverse Reinforcement Learning (offline MaxEnt-IRL) in ma…
A Lecture Note on Offline RL and IRL, Part II: Foundations of Inverse Reinforcement Learning and Dynamic Discrete Choice Models
Enoch Hyunwook Kang
In the forward reinforcement-learning problem, the reward is fixed and known; the learner is asked to find a good policy or value function. Here we turn the question around. Given…
Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity
Enoch Hyunwook Kang
Personalized alignment aims to adapt large language models to heterogeneous user preferences, yet the precise theoretical conditions for its statistical efficiency have not been fo…
Reasonably reasoning AI agents can avoid game-theoretic failures in zero-shot, provably
Enoch Hyunwook Kang
As autonomous AI agents increasingly mediate online platform markets, a fundamental question emerges: do these markets generate stable strategic outcomes? In repeated strategic env…
Demystifying the unreasonable effectiveness of online alignment methods
Enoch Hyunwook Kang
Iterative alignment methods based on purely greedy updates are remarkably effective in practice, yet existing theoretical guarantees of \(O(\log T)\) KL-regularized regret can seem…
LLM Personas as a Substitute for Field Experiments in Method Benchmarking
Enoch Hyunwook Kang
Field experiments (A/B tests) are often the most credible benchmark for methods (algorithms) in societal systems, but their cost and latency bottleneck rapid methodological progres…