1 paper · 1 filter
Ke He, Le He, Shunpu Tang +2
Expressive generative policies such as diffusion and flow models are appealing for MaxEnt online reinforcement learning because of their ability to model multimodal and highly non-…