5 citations · 6 across the 4 of their papers we have counts for
4 papers
Train Once, Get a Family: State-Adaptive Balances for Offline-to-Online Reinforcement Learning
Shenzhi Wang, Qisen Yang, Jiawei Gao +6
Offline-to-online reinforcement learning (RL) is a training paradigm that combines pre-training on a pre-collected dataset with fine-tuning in an online environment. However, the i…
Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation
Shenzhi Wang, Chang Liu, Zilong Zheng +7
Recent breakthroughs in large language models (LLMs) have brought remarkable success in the field of LLM-as-Agent. Nevertheless, a prevalent assumption is that the information proc…
Hundreds Guide Millions: Adaptive Offline Reinforcement Learning with Expert Guidance
Qisen Yang, Shenzhi Wang, Qihang Zhang +2
Offline reinforcement learning (RL) optimizes the policy on a previously collected dataset without any interactions with the environment, yet usually suffers from the distributiona…
Boosting Offline Reinforcement Learning with Action Preference Query
Qisen Yang, Shenzhi Wang, Matthieu Gaetan Lin +2
Training practical agents usually involve offline and online reinforcement learning (RL) to balance the policy's performance and interaction costs. In particular, online fine-tunin…