2 papers
cs.LG2026
Bilevel Optimization over Saddle Points of Zero-Sum Markov Games
Zihao Zheng, Irwin King, Songtao Lu
Reinforcement learning (RL) often has a hierarchical structure, where an upper-level (UL) learner selects model parameters and a lower-level (LL) decision-making process responds,…
cs.LG2025
Q-function Decomposition with Intervention Semantics with Factored Action Spaces
Junkyu Lee, Tian Gao, Elliot Nelson +3
Many practical reinforcement learning environments have a discrete factored action space that induces a large combinatorial set of actions, thereby posing significant challenges. E…