1 paper
Jihoon Park, Seungeun Oh, Seong-Lyun Kim
To address the growing demand for on-device LLM inference in resource-constrained environments, hybrid language models (HLM) have emerged, combining lightweight local models with p…