1 paper · 1 filter
Yi Hu, Cai Zhou, Muhan Zhang
The scaling of large language models (LLMs) emphasizes increasing depth, yet performance gains diminish with added layers. Prior work introduces the concept of "effective depth", a…