1 paper
Jaewoo Yang, Hayun Kim, Younghoon Kim
Modern large language models (LLMs) have established state-of-the-art performance through architectural improvements, but still require significant computational cost for inference…