1 paper
Sanghyeon Lee, Hongbeen Kim, Soojin Hwang +3
Recent large language models (LLMs) with enormous model sizes use many GPUs to meet memory capacity requirements incurring substantial costs for token generation. To provide cost-e…