7 papers
A Memory Efficient Unified Algorithm for Online Learning of Linear Dynamical Systems
Yuval Ran-Milo, Angelos Assos, Elad Hazan
Motivated by the challenge of stabilizing a general unknown linear dynamical system (LDS) from observations, we study the natural prerequisite of online prediction. Our goal is to…
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
Yuval Ran-Milo, Yotam Alexander, Shahar Mendel +1
Transformers trained via Reinforcement Learning (RL) with outcome-based supervision can spontaneously develop the ability to generate intermediate reasoning steps (Chain-of-Thought…
Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
Yuval Ran-Milo
Transformers often display an attention sink: probability mass concentrates on a fixed, content-agnostic position. Are sinks a byproduct of the optimization/training regime? Or are…
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
Yuval Ran-Milo, Hila Ofek, Shahar Mendel
Transformers commonly exhibit an attention sink: disproportionately high attention to the first position. We study this behavior in GPT-2-style models with learned query biases and…
Do Neural Networks Need Gradient Descent to Generalize? A Theoretical Study
Yotam Alexander, Yonatan Slutzky, Yuval Ran-Milo +1
Conventional wisdom attributes the mysterious generalization abilities of overparameterized neural networks to gradient descent (and its variants). The recent volume hypothesis cha…
Mamba Knockout for Unraveling Factual Information Flow
Nir Endy, Idan Daniel Grosbard, Yuval Ran-Milo +3
This paper investigates the flow of factual information in Mamba State-Space Model (SSM)-based language models. We rely on theoretical and empirical connections to Transformer-base…