1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Yen-Chen Wu, Feng-Ting Liao, Meng-Hsi Chen +3
Transformers, the standard implementation for large language models (LLMs), typically consist of tens to hundreds of discrete layers. While more layers can lead to better performan…