5 papers
Estimating time spent on work tasks
Stephane Hatgis-Kessell, Tomás Aguirre, Alexander Wan +1
The task-based framework in economics models occupations as bundles of tasks. It is the standard lens for understanding how technology affects work: a new technology changes the co…
Economic Evaluations of Language Models
Alexander Wan, Stephane Hatgis-Kessell, Tomás Aguirre +2
Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuable task. We introduce EconEvals as an ope…
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
Stephane Hatgis-Kessell, Emma Brunskill
We study when large language models (LLMs) can serve as effective black-box policy optimizers for reinforcement learning (RL) tasks, i.e., when can we replace classical RL algorith…
Influencing Humans to Conform to Preference Models for RLHF
Stephane Hatgis-Kessell, W. Bradley Knox, Serena Booth +1
Designing a reinforcement learning from human feedback (RLHF) algorithm to approximate a human's unobservable reward function requires assuming, implicitly or explicitly, a model o…
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
Stephane Hatgis-Kessell, Logan Mondal Bhamidipaty, Emma Brunskill
Human-designed reward functions for reinforcement learning (RL) agents are frequently misaligned with the humans' true, unobservable objectives, and thus act only as proxies. Optim…