1 paper
Zixuan Lan, Yanhong Li, Jiawei Zhou
Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix multiplications. We propose Reduced Matrix…