1 paper
Ziyue Liu, Ruijie Zhang, Zhengyang Wang +7
The full-size MLPs and the projection layers in attention introduce tremendous model sizes of large language models (LLMs), consuming extensive computational resources in pre-train…