10 papers
ReRAM-aware Model Finetuning addressing I-V Non-linearity and Retention Errors
Ching-Yi Lin, Shamik Kundu, Arnab Raha +1
Traditional CPU, GPU, and NPU architectures are increasingly limited by the von Neumann bottleneck. While In-Memory Computing (IMC) using ReRAM crossbar arrays offers a high-densit…
SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference
Aradhana Mohan Parvathy, Soumendu Kumar Ghosh, Shamik Kundu +4
The rapid growth in sizes of Large language models (LLMs) results in high compute and memory costs during inference. Quantization has been a significant pathway to addressing this…
COBRA: Catastrophic Bit-flip Reliability Analysis of State-Space Models
Sanjay Das, Swastik Bhattacharya, Shamik Kundu +3
State-space models (SSMs), exemplified by the Mamba architecture, have recently emerged as state-of-the-art sequence-modeling frameworks, offering linear-time scalability together…
SafeCiM: Investigating Resilience of Hybrid Floating-Point Compute-in-Memory Deep Learning Accelerators
Swastik Bhattacharya, Sanjay Das, Anand Menon +3
Deep Neural Networks (DNNs) continue to grow in complexity with Large Language Models (LLMs) incorporating vast numbers of parameters. Handling these parameters efficiently in trad…
GenBFA: An Evolutionary Optimization Approach to Bit-Flip Attacks on LLMs
Sanjay Das, Swastik Bhattacharya, Souvik Kundu +4
Large Language Models (LLMs) have revolutionized natural language processing (NLP), excelling in tasks like text generation and summarization. However, their increasing adoption in…
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
Akshat Ramachandran, Souvik Kundu, Arnab Raha +3
Large language model (LLM) pruning with fixed N:M structured sparsity significantly limits the expressivity of the sparse model, yielding sub-optimal performance. In contrast, supp…