4 papers
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation
Amr Hegazy, Amr Alanwar, Mostafa Elhoushi
Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization---…
Guiding Giants: Lightweight Controllers for Weighted Activation Steering in LLMs
Amr Hegazy, Mostafa Elhoushi, Amr Alanwar
Controlling undesirable Large Language Model (LLM) behaviors, such as the generation of unsafe content or failing to adhere to safety guidelines, often relies on costly fine-tuning…
Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM
Everlyn Asiko Chimoto, Mostafa Elhoushi, Bruce A. Bassett
Quantization is an effective technique for reducing the storage footprint and computational costs of Large Language Models (LLMs), but it often results in performance degradation.…
any4: Learned 4-bit Numeric Representation for LLMs
Mostafa Elhoushi, Jeff Johnson
We present any4, a learned 4-bit weight quantization solution for large language models (LLMs) providing arbitrary numeric representations without requiring pre-processing of weigh…