3 citations · 7 across the 5 of their papers we have counts for
5 papers
Diverse Randomized Value Functions: A Provably Pessimistic Approach for Offline Reinforcement Learning
Xudong Yu, Chenjia Bai, Hongyi Guo +2
Offline Reinforcement Learning (RL) faces distributional shift and unreliable value estimation, especially for out-of-distribution (OOD) actions. To address this, existing uncertai…
Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards
Wei Shen, Xiaoying Zhang, Yuanshun Yao +3
Reinforcement learning from human feedback (RLHF) is the mainstream paradigm used to align large language models (LLMs) with human preferences. Yet existing RLHF heavily relies on…
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
Hongyi Guo, Zhihan Liu, Yufeng Zhang +1
Large Language Models (LLMs) harness extensive data from the Internet, storing a broad spectrum of prior knowledge. While LLMs have proven beneficial as decision-making aids, their…
Human-Instruction-Free LLM Self-Alignment with Limited Samples
Hongyi Guo, Yuanshun Yao, Wei Shen +4
Aligning large language models (LLMs) with human values is a vital task for LLM practitioners. Current alignment techniques have several limitations: (1) requiring a large amount o…
Behavior Contrastive Learning for Unsupervised Skill Discovery
Rushuai Yang, Chenjia Bai, Hongyi Guo +5
In reinforcement learning, unsupervised skill discovery aims to learn diverse skills without extrinsic rewards. Previous methods discover skills by maximizing the mutual informatio…