3 papers
cs.LG2026
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
Zhao Yang, Yuxuan Jiang, Ting-Chih Chen +18
Reinforcement learning (RL) has become central to LLM post-training, yet the methods that dominate current pipelines, PPO and GRPO, represent only a narrow slice of what RL offers.…
cs.CL2026
State Design Matters: How Representations Shape Dynamic Reasoning in Large Language Models
Annie Wong, Aske Plaat, Thomas Bäck +2
As large language models (LLMs) move from static reasoning tasks toward dynamic environments, their success depends on the ability to navigate and respond to an environment that ch…
cs.AI2025
Reasoning Capabilities of Large Language Models on Dynamic Tasks
Annie Wong, Thomas Bäck, Aske Plaat +2
Large language models excel on static benchmarks, but their ability as self-learning agents in dynamic environments remains unclear. We evaluate three prompting strategies: self-re…