4 papers
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning
Renjie Mao, Xiangxin Zhou, Lvfang Tao +7
Reinforcement learning with verifiable rewards (RLVR) has become standard for improving LLM reasoning. However, existing PPO-style trust-region mechanisms remain position-agnostic…
Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning
Jiangnan Xia, Yucheng Shi, Yu Yang +3
Reinforcement learning has become a key paradigm for eliciting reasoning abilities in large language models, where exploration is crucial for discovering effective solution traject…
Grammar-Guided Evolutionary Search for Discrete Prompt Optimisation
Muzhaffar Hazman, Minh-Khoi Pham, Shweta Soundararajan +10
Prompt engineering has proven to be a crucial step in leveraging pretrained large language models (LLMs) in solving various real-world tasks. Numerous solutions have been proposed…
UCB-driven Utility Function Search for Multi-objective Reinforcement Learning
Yucheng Shi, David Lynch, Alexandros Agapitos
In Multi-objective Reinforcement Learning (MORL) agents are tasked with optimising decision-making behaviours that trade-off between multiple, possibly conflicting, objectives. MOR…