1 paper · 1 filter
Jingyao Zhang, Jaewoo Park, Jongeun Lee +1
Large Language Model (LLM) inference requires substantial computational resources, yet CPU-based inference remains essential for democratizing AI due to the widespread availability…