1 paper
Duo Chai, Zizhen Liu, Shuhuai Wang +4
Large language models (LLMs) are highly compute- and memory-intensive, posing significant demands on high-performance GPUs. At the same time, advances in GPU technology driven by s…