2 citations · 4 across the 4 of their papers we have counts for
Showing 2023Show all
3 papers · 1 filter
cs.LG2023
The Transient Nature of Emergent In-Context Learning in Transformers
Aaditya K. Singh, Stephanie C. Y. Chan, Ted Moskovitz +3
Transformer neural networks can exhibit a surprising capacity for in-context learning (ICL) despite not being explicitly trained for it. Prior work has provided a deeper understand…
cs.LG2023★ 1 cited
Confronting Reward Model Overoptimization with Constrained RLHF
Ted Moskovitz, Aaditya K. Singh, DJ Strouse +4
Large language models are typically aligned with human preferences by optimizing (RMs) fitted to human feedback. However, human preferences are multi-facet…
cs.LG2023★ 1 cited
A State Representation for Diminishing Rewards
Ted Moskovitz, Samo Hromadka, Ahmed Touati +2
A common setting in multitask reinforcement learning (RL) demands that an agent rapidly adapt to various stationary reward functions randomly sampled from a fixed distribution. In…