25 citations · 40 across the 7 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2023★ 5 cited
Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Banghua Zhu, Hiteshi Sharma, Felipe Vieira Frujeri +4
Reinforcement learning from human feedback (RLHF) has emerged as a reliable approach to aligning large language models (LLMs) to human preferences. Among the plethora of RLHF techn…
cs.CL2023
Shattering the Agent-Environment Interface for Fine-Tuning Inclusive Language Models
Wanqiao Xu, Shi Dong, Dilip Arumugam +1
A centerpiece of the ever-popular reinforcement learning from human feedback (RLHF) approach to fine-tuning autoregressive language models is the explicit training of a reward mode…