10 citations · 12 across the 6 of their papers we have counts for
6 papers · 1 filter
Boosting LLM Reasoning via Human-Inspired Reward Shaping
Wenze Lin, Zhen Yang, Xitai Jiang +2
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for enhancing reasoning in Large Language Models (LLMs). However, existing reward formulat…
Episodic Novelty Through Temporal Distance
Yuhua Jiang, Qihan Liu, Yiqin Yang +8
Exploration in sparse reward environments remains a significant challenge in reinforcement learning, particularly in Contextual Markov Decision Processes (CMDPs), where environment…
Efficient Multi-agent Reinforcement Learning by Planning
Qihan Liu, Jianing Ye, Xiaoteng Ma +3
Multi-agent reinforcement learning (MARL) algorithms have accomplished remarkable breakthroughs in solving large-scale decision-making tasks. Nonetheless, most existing MARL algori…
What is Essential for Unseen Goal Generalization of Offline Goal-conditioned RL?
Rui Yang, Yong Lin, Xiaoteng Ma +3
Offline goal-conditioned RL (GCRL) offers a way to train general-purpose agents from fully offline datasets. In addition to being conservative within the dataset, the generalizatio…
Optimistic Curiosity Exploration and Conservative Exploitation with Linear Reward Shaping
Hao Sun, Lei Han, Rui Yang +3
In this work, we study the simple yet universally applicable case of reward shaping in value-based Deep Reinforcement Learning (DRL). We show that reward shifting in the form of th…
Offline Reinforcement Learning with Value-based Episodic Memory
Xiaoteng Ma, Yiqin Yang, Hao Hu +5
Offline reinforcement learning (RL) shows promise of applying RL to real-world problems by effectively utilizing previously collected data. Most existing offline RL algorithms use…