Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Demystifying the unreasonable effectiveness of online alignment methods
Enoch Hyunwook Kang
Iterative alignment methods based on purely greedy updates are remarkably effective in practice, yet existing theoretical guarantees of \(O(\log T)\) KL-regularized regret can seem…
cs.LG2025
Stability and Generalization for Bellman Residuals
Enoch H. Kang, Kyoungseok Jang
Offline reinforcement learning and offline inverse reinforcement learning aim to recover near-optimal value functions or reward models from a fixed batch of logged trajectories, ye…