1 paper · 1 filter
Wooseong Cho, Taehyun Hwang, Joongkyu Lee +1
We study reinforcement learning with multinomial logistic (MNL) function approximation where the underlying transition probability kernel of the Markov decision processes (MDPs) is…