3 papers
cs.LG2026
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
Zhao Yang, Yuxuan Jiang, Ting-Chih Chen +18
Reinforcement learning (RL) has become central to LLM post-training, yet the methods that dominate current pipelines, PPO and GRPO, represent only a narrow slice of what RL offers.…
cs.LG2025
Aggregation of Dependent Expert Distributions in Multimodal Variational Autoencoders
Rogelio A Mancisidor, Robert Jenssen, Shujian Yu +1
Multimodal learning with variational autoencoders (VAEs) requires estimating joint distributions to evaluate the evidence lower bound (ELBO). Current methods, the product and mixtu…
cs.LG2024
ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
Hongming Li, Shujian Yu, Bin Liu +1
This paper proposes \emph{Episodic and Lifelong Exploration via Maximum ENTropy} (ELEMENT), a novel, multiscale, intrinsically motivated reinforcement learning (RL) framework that…