activity
20232025
most citedWhat needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation

2 citations · 4 across the 4 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2025

Strategy Coopetition Explains the Emergence and Transience of In-Context Learning

Aaditya K. Singh, Ted Moskovitz, Sara Dragutinovic +3

In-context learning (ICL) is a powerful ability that emerges in transformer models, enabling them to learn from context without weight updates. Recent work has established emergent…

cs.LG2024

HARP: A challenging human-annotated math reasoning benchmark

Albert S. Yue, Lovish Madaan, Ted Moskovitz +2

Math reasoning is becoming an ever increasing area of focus as we scale large language models. However, even the previously-toughest evals like MATH are now close to saturated by f…

cs.LG20242 cited

What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation

Aaditya K. Singh, Ted Moskovitz, Felix Hill +2

In-context learning is a powerful emergent ability in transformer models. Prior work in mechanistic interpretability has identified a circuit element that may be critical for in-co…

cs.LG2023

The Transient Nature of Emergent In-Context Learning in Transformers

Aaditya K. Singh, Stephanie C. Y. Chan, Ted Moskovitz +3

Transformer neural networks can exhibit a surprising capacity for in-context learning (ICL) despite not being explicitly trained for it. Prior work has provided a deeper understand…

cs.LG20231 cited

Confronting Reward Model Overoptimization with Constrained RLHF

Ted Moskovitz, Aaditya K. Singh, DJ Strouse +4

Large language models are typically aligned with human preferences by optimizing (RMs) fitted to human feedback. However, human preferences are multi-facet…

cs.LG20231 cited

A State Representation for Diminishing Rewards

Ted Moskovitz, Samo Hromadka, Ahmed Touati +2

A common setting in multitask reinforcement learning (RL) demands that an agent rapidly adapt to various stationary reward functions randomly sampled from a fixed distribution. In…