Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
ArcLight: A Lightweight LLM Inference Architecture for Many-Core CPUs
Yuzhuang Xu, Xu Han, Yuxuan Li +1
Although existing frameworks for large language model (LLM) inference on CPUs are mature, they fail to fully exploit the computation potential of many-core CPU platforms. Many-core…
cs.DC2024
Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices
Xueyuan Han, Zinuo Cai, Yichu Zhang +4
The application of Transformer-based large models has achieved numerous success in recent years. However, the exponential growth in the parameters of large models introduces formid…