2 papers
cs.AI2024
The Asymptotic Behavior of Attention in Transformers
Álvaro Rodríguez Abella, João Pedro Silvestre, Paulo Tabuada
The transformer architecture has become the foundation of modern Large Language Models (LLMs), yet its theoretical properties are still not well understood. As with classic neural…
math.OC2024
A Framework for Time-Varying Optimization via Derivative Estimation
Matteo Marchi, Jonathan Bunton, João Pedro Silvestre +1
Optimization algorithms have a rich and fundamental relationship with ordinary differential equations given by its continuous-time limit. When the cost function varies with time --…