4 papers
MoKA: Mixture of Kronecker Adapters
Mohammadreza Sadeghi, Mahsa Ghazvini Nejad, MirHamed Jafarzadeh Asl +4
Parameter-efficient fine-tuning (PEFT) is essential for reducing the computational overhead of large language models (LLMs). Low-rank family adapters are commonly used to control t…
Tiny Noise-Robust Voice Activity Detector for Voice Assistants
Hamed Jafarzadeh Asl, Mahsa Ghazvini Nejad, Amin Edraki +2
Voice Activity Detection (VAD) in the presence of background noise remains a challenging problem in speech processing. Accurate VAD is essential in automatic speech recognition, vo…
OAC: Output-adaptive Calibration for Accurate Post-training Quantization
Ali Edalati, Alireza Ghaffari, Mahsa Ghazvini Nejad +4
Deployment of Large Language Models (LLMs) has major computational costs, due to their rapidly expanding size. Compression of LLMs reduces the memory footprint, latency, and energy…
Rethinking Post-Training Quantization: Introducing a Statistical Pre-Calibration Approach
Alireza Ghaffari, Sharareh Younesian, Boxing Chen +2
As Large Language Models (LLMs) become increasingly computationally complex, developing efficient deployment strategies, such as quantization, becomes crucial. State-of-the-art Pos…