1 paper
Ke He, Le He, Shunpu Tang +2
Expressive generative policies such as diffusion and flow models are appealing for MaxEnt online reinforcement learning because of their ability to model multimodal and highly non-…