Publications (6)
Performance Optimization of Ratings-Based Reinforcement Learning
Evelyn Rose, Devin White, Mingkang Wu +3
This paper explores multiple optimization methods to improve the performance of rating-based reinforcement learning (RbRL). RbRL, a method based on the idea of human ratings, has b…
Atari-GPT: Benchmarking Multimodal Large Language Models as Low-Level Policies in Atari Games
Nicholas R. Waytowich, Devin White, MD Sunbeam +1
Recent advancements in large language models (LLMs) have expanded their capabilities beyond traditional text-based tasks to multimodal domains, integrating visual, auditory, and te…
Human-Inspired Multi-Level Reinforcement Learning
Mingkang Wu, Devin White, Vernon Lawhern +2
Reinforcement learning (RL), a common tool in decision making, learns control policies from various experiences based on the associated cumulative return/rewards without treating t…
Rating-based Reinforcement Learning
Devin White, Mingkang Wu, Ellen Novoseller +3
This paper develops a novel rating-based reinforcement learning approach that uses human ratings to obtain human guidance in reinforcement learning. Different from the existing pre…
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
Joshua Barron, Devin White
The relationship between memorization and generalization in large language models (LLMs) remains an open area of research, with growing evidence that the two are deeply intertwined…
Multi-Task Reward Learning from Human Ratings
Mingkang Wu, Devin White, Evelyn Rose +3
Reinforcement learning from human feedback (RLHF) has become a key factor in aligning model behavior with users' goals. However, while humans integrate multiple strategies when mak…