1 paper
Wooseong Cho, Taehyun Hwang, Joongkyu Lee +1
We study reinforcement learning with multinomial logistic (MNL) function approximation where the underlying transition probability kernel of the Markov decision processes (MDPs) is…