1 paper · 1 filter
Andrei Baroian, Kasper Notebomer
Transformer-based language models traditionally use uniform (isotropic) layer sizes, yet they ignore the diverse functional roles that different depths can play and their computati…