1 paper · 1 filter
Jay H. Park, Youngju Cho, Choungsol Lee +2
Large language model (LLM) inference often suffers from high latency, particularly in resource-constrained environments such as on-device or edge deployments. To address this chall…