papers

Publications (6)

cs.LG2025

Performance Optimization of Ratings-Based Reinforcement Learning

Evelyn Rose, Devin White, Mingkang Wu +3

This paper explores multiple optimization methods to improve the performance of rating-based reinforcement learning (RbRL). RbRL, a method based on the idea of human ratings, has b…

cs.AI2024

Atari-GPT: Benchmarking Multimodal Large Language Models as Low-Level Policies in Atari Games

Nicholas R. Waytowich, Devin White, MD Sunbeam +1

Recent advancements in large language models (LLMs) have expanded their capabilities beyond traditional text-based tasks to multimodal domains, integrating visual, auditory, and te…

cs.LG2025

Human-Inspired Multi-Level Reinforcement Learning

Mingkang Wu, Devin White, Vernon Lawhern +2

Reinforcement learning (RL), a common tool in decision making, learns control policies from various experiences based on the associated cumulative return/rewards without treating t…

cs.LG2024

Rating-based Reinforcement Learning

Devin White, Mingkang Wu, Ellen Novoseller +3

This paper develops a novel rating-based reinforcement learning approach that uses human ratings to obtain human guidance in reinforcement learning. Different from the existing pre…

cs.LG2025

Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers

Joshua Barron, Devin White

The relationship between memorization and generalization in large language models (LLMs) remains an open area of research, with growing evidence that the two are deeply intertwined…

cs.LG2025

Multi-Task Reward Learning from Human Ratings

Mingkang Wu, Devin White, Evelyn Rose +3

Reinforcement learning from human feedback (RLHF) has become a key factor in aligning model behavior with users' goals. However, while humans integrate multiple strategies when mak…