5 citations · 5 across the 3 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
Anna Hedström, Salim I. Amoukou, Tom Bewley +2
We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventi…
cs.LG2025
Zero-Shot Reinforcement Learning Under Partial Observability
Scott Jeen, Tom Bewley, Jonathan M. Cullen
Recent work has shown that, under certain assumptions, zero-shot reinforcement learning (RL) methods can generalise to any unseen task in an environment after reward-free pre-train…
cs.LG2021★ 5 cited
Interpretable Preference-based Reinforcement Learning with Tree-Structured Reward Functions
Tom Bewley, Freddy Lecue
The potential of reinforcement learning (RL) to deliver aligned and performant agents is partially bottlenecked by the reward engineering problem. One alternative to heuristic tria…