1 paper
Fangxin Liu, Qinghua Zhang, Hanjing Shen +5
The rapid evolution of Large Language Models (LLMs) towards long-context reasoning and sparse architectures has pushed memory requirements far beyond the capacity of individual dev…