5 papers
Redistribution-based Cost Inference Improves Sparse Safe Offline RL
Ebenezer Gelo, Geraud Nangue Tasse, Steven James +1
Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first u…
CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models
Siddarth Singh, Victoria Williams, Simon Rosen +6
The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict. Current evaluati…
Position: the Stochastic Parrot in the Coal Mine. Model Collapse is a Threat to Low-Resource Communities
Devon Jarvis, Richard Klein, Benjamin Rosman +2
Model collapse, the degradation in performance that arises when generative models are trained on the outputs of prior models, is an increasing concern as artificially generated con…
Unsupervised Hierarchical Skill Discovery
Damion Harvey, Geraud Nangue Tasse, Benjamin Rosman +2
We consider the problem of unsupervised skill segmentation and hierarchical structure discovery in reinforcement learning. While recent approaches have sought to segment trajectori…
Compositional Instruction Following with Language Models and Reinforcement Learning
Vanya Cohen, Geraud Nangue Tasse, Nakul Gopalan +4
Combining reinforcement learning with language grounding is challenging as the agent needs to explore the environment while simultaneously learning multiple language-conditioned ta…