5 citations · 9 across the 4 of their papers we have counts for
4 papers
What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
Aaditya K. Singh, Ted Moskovitz, Felix Hill +2
In-context learning is a powerful emergent ability in transformer models. Prior work in mechanistic interpretability has identified a circuit element that may be critical for in-co…
Confronting Reward Model Overoptimization with Constrained RLHF
Ted Moskovitz, Aaditya K. Singh, DJ Strouse +4
Large language models are typically aligned with human preferences by optimizing (RMs) fitted to human feedback. However, human preferences are multi-facet…
A State Representation for Diminishing Rewards
Ted Moskovitz, Samo Hromadka, Ahmed Touati +2
A common setting in multitask reinforcement learning (RL) demands that an agent rapidly adapt to various stationary reward functions randomly sampled from a fixed distribution. In…
ReLOAD: Reinforcement Learning with Optimistic Ascent-Descent for Last-Iterate Convergence in Constrained MDPs
Ted Moskovitz, Brendan O'Donoghue, Vivek Veeriah +3
In recent years, Reinforcement Learning (RL) has been applied to real-world problems with increasing success. Such applications often require to put constraints on the agent's beha…