Showing 2024Show all
3 papers · 1 filter
cs.LG2024
Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of Transformers
Lorenzo Tiberi, Francesca Mignacco, Kazuki Irie +1
Despite the remarkable empirical performance of Transformers, their theoretical understanding remains elusive. Here, we consider a deep multi-head self-attention network, that is c…
cs.LG2024
Diverse capability and scaling of diffusion and auto-regressive models when learning abstract rules
Binxu Wang, Jiaqi Shang, Haim Sompolinsky
Humans excel at discovering regular structures from limited samples and applying inferred rules to novel settings. We investigate whether modern generative models can similarly lea…
cs.LG2024
Coding schemes in neural networks learning classification tasks
Alexander van Meegen, Haim Sompolinsky
Neural networks posses the crucial ability to generate meaningful representations of task-dependent features. Indeed, with appropriate scaling, supervised learning in neural networ…