1 paper · 1 filter
Sergii Kozyrev, Davyd Maiboroda
Large language models are limited in deployment by GPU memory and inference latency. We present Minima, a production compression pipeline that learns where and how to structurally…