2 papers
cs.LG2024
Limits of Transformer Language Models on Learning to Compose Algorithms
Jonathan Thomm, Giacomo Camposampiero, Aleksandar Terzic +3
We analyze the capabilities of Transformer language models in learning compositional discrete tasks. To this end, we evaluate training LLaMA models and prompting GPT-4 and Gemini o…
cs.LG2023
TCNCA: Temporal Convolution Network with Chunked Attention for Scalable Sequence Processing
Aleksandar Terzic, Michael Hersche, Geethan Karunaratne +3
MEGA is a recent transformer-based architecture, which utilizes a linear recurrent operator whose parallel computation, based on the FFT, scales as , with being the s…