6 papers · 1 filter
Preference-Based Reward Learning under Partial Observability with Inexact Dynamics
Reza Zolnouri, Semih Cayci
In this paper, we study how partial observability and inexact latent-state inference affect reward learning from preferences. To that end, we study preference-based reward learning…
A Riemannian Optimization Perspective of the Gauss-Newton Method for Feedforward Neural Networks
Semih Cayci
In this work, we establish non-asymptotic convergence bounds for the Gauss-Newton method in training neural networks with smooth activations. In the underparameterized regime, the…
Optimal Rates of Convergence for Entropy Regularization in Discounted Markov Decision Processes
Johannes Müller, Semih Cayci
We study the error introduced by entropy regularization in infinite-horizon discrete discounted Markov decision processes. We show that this error decreases exponentially in the in…
Recurrent Natural Policy Gradient for POMDPs
Semih Cayci, Atilla Eryilmaz
Solving partially observable Markov decision processes (POMDPs) remains a fundamental challenge in reinforcement learning (RL), primarily due to the curse of dimensionality induced…
Fisher-Rao Gradient Flows of Linear Programs and State-Action Natural Policy Gradients
Johannes Müller, Semih Ãaycı, Guido Montúfar
Kakade's natural policy gradient method has been studied extensively in recent years, showing linear convergence with and without regularization. We study another natural gradient…
A Variational Inequality Approach to Independent Learning in Static Mean-Field Games
Batuhan Yardim, Semih Cayci, Niao He
Competitive games involving thousands or even millions of players are prevalent in real-world contexts, such as transportation, communications, and computer networks. However, lear…