4 papers
EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models
Jinane Bazzi, Mariam Rakka, Fadi Kurdahi +2
The growing demand for deploying Small Language Models (SLMs) on edge devices, including laptops, smartphones, and embedded platforms, has exposed fundamental inefficiencies in exi…
Mixed-Precision Quantization for Language Models: Techniques and Prospects
Mariam Rakka, Marios Fournarakis, Olga Krestinskaya +5
The rapid scaling of language models (LMs) has resulted in unprecedented computational, memory, and energy requirements, making their training and deployment increasingly unsustain…
SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
Mariam Rakka, Jinhao Li, Guohao Dai +3
Recent research efforts focus on reducing the computational and memory overheads of Large Language Models (LLMs) to make them feasible on resource-constrained devices. Despite adva…
BF-IMNA: A Bit Fluid In-Memory Neural Architecture for Neural Network Acceleration
Mariam Rakka, Rachid Karami, Ahmed M. Eltawil +2
Mixed-precision quantization works Neural Networks (NNs) are gaining traction for their efficient realization on the hardware leading to higher throughput and lower energy. In-Memo…