collaborators

6 papers

cs.LG2026

IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression

Ali Abbasi, Chayne Thrash, Haoran Qin +2

Large language models deliver strong performance across language and reasoning tasks, but their storage and compute costs remain major barriers to deployment in resource-constraine…

cs.LG2026

ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs

Chayne Thrash, Ali Abbasi, Soheil Kolouri

Large language models (LLMs) are costly to deploy due to their large memory footprint and high inference cost. Weight-activation quantization can reduce these costs, but low-bit ac…

cs.LG2026

LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport

Ashkan Shahbazi, Chayne Thrash, Yikun Bai +3

Transformers have proven highly effective across modalities, but standard softmax attention scales quadratically with sequence length, limiting long context modeling. Linear attent…

cs.LG2026

Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression

Ali Abbasi, Chayne Thrash, Haoran Qin +3

Advances in large language models have driven strong performance across many tasks, but their memory and compute costs still hinder deployment. SVD-based compression reduces storag…

cs.LG2025

Low-Rank Prehab: Preparing Neural Networks for SVD Compression

Haoran Qin, Shansita Sharma, Ali Abbasi +2

Low-rank approximation methods such as singular value decomposition (SVD) and its variants (e.g., Fisher-weighted SVD, Activation SVD) have recently emerged as effective tools for…

cs.LG2025

MCNC: Manifold-Constrained Reparameterization for Neural Compression

Chayne Thrash, Ali Abbasi, Reed Andreas +4

The outstanding performance of large foundational models across diverse tasks, from computer vision to speech and natural language processing, has significantly increased their dem…