Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
The Asymptotic Behavior of Attention in Transformers
Ãlvaro RodrÃguez Abella, João Pedro Silvestre, Paulo Tabuada
The transformer architecture has become the foundation of modern Large Language Models (LLMs), yet its theoretical properties are still not well understood. As with classic neural…
cs.AI2024
Meanings and Feelings of Large Language Models: Observability of Latent States in Generative AI
Tian Yu Liu, Stefano Soatto, Matteo Marchi +2
We tackle the question of whether Large Language Models (LLMs), viewed as dynamical systems with state evolving in the embedding space of symbolic tokens, are observable. That is,…