4 papers
Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models
Deepanshu Pandey, Arnav Chavan, Nahush Lele +2
Quantization of Large Language Models (LLMs) is often hindered by the sensitivity of the self-attention mechanism to discretization errors. We identify the softmax operator as a bo…
S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations
Arnav Chavan, Nahush Lele, Udbhav Bamba +3
Activation outliers in large-scale transformer models pose a fundamental challenge to model quantization, creating excessively large ranges that cause severe accuracy drops during…
Surgical Feature-Space Decomposition of LLMs: Why, When and How?
Arnav Chavan, Nahush Lele, Deepak Gupta
Low-rank approximations, of the weight and feature space can enhance the performance of deep learning models, whether in terms of improving generalization or reducing the latency o…
Rethinking Compression: Reduced Order Modelling of Latent Features in Large Language Models
Arnav Chavan, Nahush Lele, Deepak Gupta
Due to the substantial scale of Large Language Models (LLMs), the direct application of conventional compression methodologies proves impractical. The computational demands associa…