1 paper
Bodon Jeong, Hongsu Byun, Youngjae Kim +4
The increasing deployment of Large Language Model (LLM) inference on edge AI systems demands efficient execution under tight memory budgets. A key challenge arises from Key-Value (…