Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Modeling Transformers as complex networks to analyze learning dynamics
Elisabetta Rocchetti
The process by which Large Language Models (LLMs) acquire complex capabilities during training remains a key open question in mechanistic interpretability. This project investigate…
cs.LG2024
Unveiling Transformer Perception by Exploring Input Manifolds
Alessandro Benfenati, Alfio Ferrara, Alessio Marta +2
This paper introduces a general method for the exploration of equivalence classes in the input space of Transformer models. The proposed approach is based on sound mathematical the…