1 paper
Theophilus Amaefuna, Hitesh Vaidya, Anshuman Chhabra +1
Layer-wise capacity in large language models is highly non-uniform: some layers contribute disproportionately to loss reduction, whereas others are nearly redundant. Existing layer…