2 papers
cs.AR2026
BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators
Fabian Waschkowski, Prabod Rathnayaka, Lukas Wesemann
Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor AP…
cs.CL2026
BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal
Prabod Rathnayaka, Fabian Waschkowski, Lukas Wesemann
We present BaseRT, a native Metal inference runtime for large language models (LLMs) on Apple Silicon, and report the highest inference throughput on this hardware to date. Existin…