1 paper · 1 filter
Hung-Yueh Chiang, Bokun Wang, Diana Marculescu
The latency and power consumption of large language models (LLMs) are major constraints when serving them across a wide spectrum of hardware platforms, from mobile edge devices to…