Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers
Takuya Ito, Ruchir Puri, Murray Campbell +1
Learning generalizable algorithmic computations remains a challenge for neural networks, as reflected in persistent failures on compositional and length generalization benchmarks.…
cs.LG2024
Learning interpretable positional encodings in transformers depends on initialization
Takuya Ito, Luca Cocchi, Tim Klinger +3
In transformers, the positional encoding (PE) provides essential information that distinguishes the position and order amongst tokens in a sequence. Most prior investigations of PE…
cs.LG2024
On the generalization capacity of neural networks during generic multimodal reasoning
Takuya Ito, Soham Dan, Mattia Rigotti +2
The advent of the Transformer has led to the development of large language models (LLM), which appear to demonstrate human-like capabilities. To assess the generality of this class…