3 papers
cs.LG2024
Transformers on Markov Data: Constant Depth Suffices
Nived Rajaraman, Marco Bondaschi, Kannan Ramchandran +2
Attention-based transformers have been remarkably successful at modeling generative processes across various domains and modalities. In this paper, we study the behavior of transfo…
cs.LG2024
Local to Global: Learning Dynamics and Effect of Initialization for Transformers
Ashok Vardhan Makkuva, Marco Bondaschi, Chanakya Ekbote +4
In recent years, transformer-based models have revolutionized deep learning, particularly in sequence modeling. To better understand this phenomenon, there is a growing interest in…
cs.IT2024
Batch Universal Prediction
Marco Bondaschi, Michael Gastpar
Large language models (LLMs) have recently gained much popularity due to their surprising ability at generating human-like English sentences. LLMs are essentially predictors, estim…