4 papers
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning
Zhao Yang, Yuxuan Jiang, Ting-Chih Chen +18
Reinforcement learning (RL) has become central to LLM post-training, yet the methods that dominate current pipelines, PPO and GRPO, represent only a narrow slice of what RL offers.…
State Design Matters: How Representations Shape Dynamic Reasoning in Large Language Models
Annie Wong, Aske Plaat, Thomas Bäck +2
As large language models (LLMs) move from static reasoning tasks toward dynamic environments, their success depends on the ability to navigate and respond to an environment that ch…
Multi-Step Reasoning with Large Language Models, a Survey
Aske Plaat, Annie Wong, Suzan Verberne +3
Large language models (LLMs) with billions of parameters exhibit in-context learning abilities, enabling few-shot learning on tasks that the model was not specifically trained for.…
Reasoning Capabilities of Large Language Models on Dynamic Tasks
Annie Wong, Thomas Bäck, Aske Plaat +2
Large language models excel on static benchmarks, but their ability as self-learning agents in dynamic environments remains unclear. We evaluate three prompting strategies: self-re…