Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Transformer Circuits Can Realize Clustering Algorithms
Kenneth L. Clarkson, Lior Horesh, Takuya Ito +2
Although transformers are most commonly optimized as statistical sequence models, it is unclear to what extent they can implement and learn exact algorithmic computations. Here, we…
cs.LG2025
Learning interpretable positional encodings in transformers depends on initialization
Takuya Ito, Luca Cocchi, Tim Klinger +3
In transformers, the positional encoding (PE) provides essential information that distinguishes the position and order amongst tokens in a sequence. Most prior investigations of PE…