1 paper
Xiangfeng Wang, Zaiyi Chen, Zheyong Xie +3
With the rising popularity of Transformer-based large language models (LLMs), reducing their high inference costs has become a significant research focus. One effective approach is…