226 citations · 228 across the 3 of their papers we have counts for
3 papers
cs.LG2019
Policy Optimization Through Approximate Importance Sampling
Marcin B. Tomczak, Dongho Kim, Peter Vrancx +1
Recent policy optimization approaches (Schulman et al., 2015a; 2017) have achieved substantial empirical successes by constructing new proxy optimization objectives. These proxy ob…
cs.AI2014★ 226 cited
Learning to Cooperate via Policy Search
Leonid Peshkin, Kee-Eung Kim, Nicolas Meuleau +1
Cooperative games are those in which both agents share the same payoff structure. Value-based reinforcement-learning algorithms, such as variants of Q-learning, have been applied t…
cs.AI2012★ 2 cited
A Geometric Traversal Algorithm for Reward-Uncertain MDPs
Eunsoo Oh, Kee-Eung Kim
Markov decision processes (MDPs) are widely used in modeling decision making problems in stochastic environments. However, precise specification of the reward functions in MDPs is…