collaborators

5 papers

cs.LG2026

Redistribution-based Cost Inference Improves Sparse Safe Offline RL

Ebenezer Gelo, Geraud Nangue Tasse, Steven James +1

Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first u…

cs.AI2026

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

Siddarth Singh, Victoria Williams, Simon Rosen +6

The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict. Current evaluati…

cs.LG2026

Position: the Stochastic Parrot in the Coal Mine. Model Collapse is a Threat to Low-Resource Communities

Devon Jarvis, Richard Klein, Benjamin Rosman +2

Model collapse, the degradation in performance that arises when generative models are trained on the outputs of prior models, is an increasing concern as artificially generated con…

cs.LG2026

Unsupervised Hierarchical Skill Discovery

Damion Harvey, Geraud Nangue Tasse, Benjamin Rosman +2

We consider the problem of unsupervised skill segmentation and hierarchical structure discovery in reinforcement learning. While recent approaches have sought to segment trajectori…

cs.LG2025

Compositional Instruction Following with Language Models and Reinforcement Learning

Vanya Cohen, Geraud Nangue Tasse, Nakul Gopalan +4

Combining reinforcement learning with language grounding is challenging as the agent needs to explore the environment while simultaneously learning multiple language-conditioned ta…