7 citations · 8 across the 2 of their papers we have counts for
2 papers
cs.LG2023★ 1 cited
Effective Theory of Transformers at Initialization
Emily Dinan, Sho Yaida, Susan Zhang
We perform an effective-theory analysis of forward-backward signal propagation in wide and deep Transformers, i.e., residual neural networks with multi-head self-attention blocks a…
cs.CL2023★ 7 cited
Scaling Laws for Generative Mixed-Modal Language Models
Armen Aghajanyan, Lili Yu, Alexis Conneau +7
Generative language models define distributions over sequences of tokens that can represent essentially any combination of data modalities (e.g., any permutation of image tokens fr…