6 papers
IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression
Ali Abbasi, Chayne Thrash, Haoran Qin +2
Large language models deliver strong performance across language and reasoning tasks, but their storage and compute costs remain major barriers to deployment in resource-constraine…
ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs
Chayne Thrash, Ali Abbasi, Soheil Kolouri
Large language models (LLMs) are costly to deploy due to their large memory footprint and high inference cost. Weight-activation quantization can reduce these costs, but low-bit ac…
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
Ashkan Shahbazi, Chayne Thrash, Yikun Bai +3
Transformers have proven highly effective across modalities, but standard softmax attention scales quadratically with sequence length, limiting long context modeling. Linear attent…
Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
Ali Abbasi, Chayne Thrash, Haoran Qin +3
Advances in large language models have driven strong performance across many tasks, but their memory and compute costs still hinder deployment. SVD-based compression reduces storag…
Low-Rank Prehab: Preparing Neural Networks for SVD Compression
Haoran Qin, Shansita Sharma, Ali Abbasi +2
Low-rank approximation methods such as singular value decomposition (SVD) and its variants (e.g., Fisher-weighted SVD, Activation SVD) have recently emerged as effective tools for…
MCNC: Manifold-Constrained Reparameterization for Neural Compression
Chayne Thrash, Ali Abbasi, Reed Andreas +4
The outstanding performance of large foundational models across diverse tasks, from computer vision to speech and natural language processing, has significantly increased their dem…