3 papers
cs.LG2024
On the consistency of hyper-parameter selection in value-based deep reinforcement learning
Johan Obando-Ceron, João G. M. Araújo, Aaron Courville +1
Deep reinforcement learning (deep RL) has achieved tremendous success on various domains through a combination of algorithmic design and careful selection of hyper-parameters. Algo…
cs.CL2024
Transformers need glasses! Information over-squashing in language tasks
Federico Barbero, Andrea Banino, Steven Kapturowski +5
We study how information propagates in decoder-only Transformers, which are the architectural backbone of most existing frontier large language models (LLMs). We rely on a theoreti…
cs.LG2024
Position: Categorical Deep Learning is an Algebraic Theory of All Architectures
Bruno GavranoviÄ, Paul Lessard, Andrew Dudzik +3
We present our position on the elusive quest for a general-purpose framework for specifying and studying deep learning architectures. Our opinion is that the key attempts made so f…