36 citations · 39 across the 15 of their papers we have counts for
6 papers · 2 filters
Batch Normalization Decomposed
Ido Nachum, Marco Bondaschi, Michael Gastpar +1
\emph{Batch normalization} is a successful building block of neural network architectures. Yet, it is not well understood. A neural network layer with batch normalization comprises…
Which Algorithms Have Tight Generalization Bounds?
Michael Gastpar, Ido Nachum, Jonathan Shafer +1
We study which machine learning algorithms have tight generalization bounds. First, we present conditions that preclude the existence of tight generalization bounds. Specifically,…
Transformers on Markov Data: Constant Depth Suffices
Nived Rajaraman, Marco Bondaschi, Kannan Ramchandran +2
Attention-based transformers have been remarkably successful at modeling generative processes across various domains and modalities. In this paper, we study the behavior of transfo…
Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models
Alliot Nagle, Adway Girish, Marco Bondaschi +3
We formalize the problem of prompt compression for large language models (LLMs) and present a framework to unify token-level prompt compression methods which create hard prompts fo…
Local to Global: Learning Dynamics and Effect of Initialization for Transformers
Ashok Vardhan Makkuva, Marco Bondaschi, Chanakya Ekbote +4
In recent years, transformer-based models have revolutionized deep learning, particularly in sequence modeling. To better understand this phenomenon, there is a growing interest in…
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
Ashok Vardhan Makkuva, Marco Bondaschi, Adway Girish +4
Attention-based transformers have achieved tremendous success across a variety of disciplines including natural languages. To deepen our understanding of their sequential modeling…