Showing 2025Show all
2 papers · 1 filter
cs.LG2025
Mixed-Precision Quantization for Language Models: Techniques and Prospects
Mariam Rakka, Marios Fournarakis, Olga Krestinskaya +5
The rapid scaling of language models (LMs) has resulted in unprecedented computational, memory, and energy requirements, making their training and deployment increasingly unsustain…
cs.AI2025
CIMNAS: A Joint Framework for Compute-In-Memory-Aware Neural Architecture Search
Olga Krestinskaya, Mohammed E. Fouda, Ahmed Eltawil +1
To maximize hardware efficiency and performance accuracy in Compute-In-Memory (CIM)-based neural network accelerators for Artificial Intelligence (AI) applications, co-optimizing b…