2 papers
cs.LG2026
A Controlled Study of Attention-Only Transformers
Henry Ndubuaku, Karen Mosoyan, Jakub Mroz +5
Feed-forward networks hold two thirds of a transformer's non-embedding parameters, yet the architecture has not received a necessity test that controls parameters, compute, and dep…
cs.CL2025
Parameter-Efficient Transformer Embeddings
Henry Ndubuaku, Mouad Talhi
Embedding layers in transformer-based NLP models typically account for the largest share of model parameters, scaling with vocabulary size but not yielding performance gains propor…