Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Forager: a lightweight testbed for continual learning with partial observability in RL
Steven Tang, Xinze Xiong, Anna Hakhverdyan +7
In continual reinforcement learning (CRL), good performance requires never-ending learning, acting, and exploration in a big, partially observable world. Most CRL experiments have…
cs.LG2024
Understanding and Alleviating Memory Consumption in RLHF for LLMs
Jin Zhou, Hanmei Yang, Steven +4
Fine-tuning with Reinforcement Learning with Human Feedback (RLHF) is essential for aligning large language models (LLMs). However, RLHF often encounters significant memory challen…