9 citations · 12 across the 3 of their papers we have counts for
3 papers
cs.LG2024★ 2 cited
What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
Aaditya K. Singh, Ted Moskovitz, Felix Hill +2
In-context learning is a powerful emergent ability in transformer models. Prior work in mechanistic interpretability has identified a circuit element that may be critical for in-co…
cs.CL2024★ 9 cited
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
Aaditya K. Singh, DJ Strouse
Tokenization, the division of input text into input tokens, is an often overlooked aspect of the large language model (LLM) pipeline and could be the source of useful or harmful in…
cs.LG2023★ 1 cited
Confronting Reward Model Overoptimization with Constrained RLHF
Ted Moskovitz, Aaditya K. Singh, DJ Strouse +4
Large language models are typically aligned with human preferences by optimizing (RMs) fitted to human feedback. However, human preferences are multi-facet…