1 paper · 1 filter
Prabod Rathnayaka, Fabian Waschkowski, Lukas Wesemann
We present BaseRT, a native Metal inference runtime for large language models (LLMs) on Apple Silicon, and report the highest inference throughput on this hardware to date. Existin…