most citedBehavior Contrastive Learning for Unsupervised Skill Discovery

3 citations · 7 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG20241 cited

Diverse Randomized Value Functions: A Provably Pessimistic Approach for Offline Reinforcement Learning

Xudong Yu, Chenjia Bai, Hongyi Guo +2

Offline Reinforcement Learning (RL) faces distributional shift and unreliable value estimation, especially for out-of-distribution (OOD) actions. To address this, existing uncertai…

cs.CL2024

Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards

Wei Shen, Xiaoying Zhang, Yuanshun Yao +3

Reinforcement learning from human feedback (RLHF) is the mainstream paradigm used to align large language models (LLMs) with human preferences. Yet existing RLHF heavily relies on…

cs.AI20241 cited

Can Large Language Models Play Games? A Case Study of A Self-Play Approach

Hongyi Guo, Zhihan Liu, Yufeng Zhang +1

Large Language Models (LLMs) harness extensive data from the Internet, storing a broad spectrum of prior knowledge. While LLMs have proven beneficial as decision-making aids, their…

cs.CL20242 cited

Human-Instruction-Free LLM Self-Alignment with Limited Samples

Hongyi Guo, Yuanshun Yao, Wei Shen +4

Aligning large language models (LLMs) with human values is a vital task for LLM practitioners. Current alignment techniques have several limitations: (1) requiring a large amount o…

cs.LG20233 cited

Behavior Contrastive Learning for Unsupervised Skill Discovery

Rushuai Yang, Chenjia Bai, Hongyi Guo +5

In reinforcement learning, unsupervised skill discovery aims to learn diverse skills without extrinsic rewards. Previous methods discover skills by maximizing the mutual informatio…