23 citations · 33 across the 3 of their papers we have counts for
4 papers
Model compression as constrained optimization, with application to neural nets. Part V: combining compressions
Miguel Á. Carreira-Perpiñán, Yerlan Idelbayev
Model compression is generally performed by using quantization, low-rank approximation or pruning, for which various algorithms have been researched in recent years. One fundamenta…
A flexible, extensible software framework for model compression based on the LC algorithm
Yerlan Idelbayev, Miguel Á. Carreira-Perpiñán
We propose a software framework based on the ideas of the Learning-Compression (LC) algorithm, that allows a user to compress a neural network or other machine learning model using…
Structured Multi-Hashing for Model Compression
Elad Eban, Yair Movshovitz-Attias, Hao Wu +4
Despite the success of deep neural networks (DNNs), state-of-the-art models are too large to deploy on low-resource devices or common server configurations in which multiple models…
Model compression as constrained optimization, with application to neural nets. Part II: quantization
Miguel Á. Carreira-Perpiñán, Yerlan Idelbayev
We consider the problem of deep neural net compression by quantization: given a large, reference net, we want to quantize its real-valued weights using a codebook with entries…