Showing cs.ARShow all
2 papers · 1 filter
cs.AR2025
Hardware-Software Co-Design for Accelerating Transformer Inference Leveraging Compute-in-Memory
Dong Eun Kim, Tanvi Sharma, Kaushik Roy
Transformers have become the backbone of neural network architecture for most machine learning applications. Their widespread use has resulted in multiple efforts on accelerating a…
cs.AR2023
WWW: What, When, Where to Compute-in-Memory
Tanvi Sharma, Mustafa Ali, Indranil Chakraborty +1
Matrix multiplication is the dominant computation during Machine Learning (ML) inference. To efficiently perform such multiplication operations, Compute-in-memory (CiM) paradigms h…