1 paper
Kaivan Kamali, Kajetan Schweighofer, Hormoz Shahrzad +3
The massive scaling of Large Language Models (LLMs) has made pretraining increasingly cost-prohibitive. While low-rank representation and orthonormal weight matrices could in princ…