1 paper
Sergii Kozyrev, Davyd Maiboroda
Large language models are limited in deployment by GPU memory and inference latency. We present Minima, a production compression pipeline that learns where and how to structurally…