Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
A Mixture-of-Experts Approach to Few-Shot Task Transfer in Open-Ended Text Worlds
Christopher Z. Cui, Xiangyu Peng, Mark O. Riedl
Open-ended worlds are those in which there are no pre-specified goals or environmental reward signal. As a consequence, an agent must know how to perform a multitude of tasks. Howe…
cs.CL2024
Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
Ashutosh Baheti, Ximing Lu, Faeze Brahman +3
Reinforcement Learning with Human Feedback (RLHF) is the most prominent method for Language Model (LM) alignment. However, RLHF is an unstable and data-hungry process that continua…