collaborators

15 papers

cs.LG2026

ReRAM-aware Model Finetuning addressing I-V Non-linearity and Retention Errors

Ching-Yi Lin, Shamik Kundu, Arnab Raha +1

Traditional CPU, GPU, and NPU architectures are increasingly limited by the von Neumann bottleneck. While In-Memory Computing (IMC) using ReRAM crossbar arrays offers a high-densit…

cs.AR2026

KATANA: A Fast, Low-Power Mapping of Kalman Filters onto Edge NPUs for Real-Time Tracking

Bodhisatwa Kundu, Anish Rooj, Sumit Saha +4

State estimation is the closed-loop core of every real-time tracking system, from radar surveillance and counter-UAV defense to autonomous driving and robotics. These deployments r…

cs.AR2026

MOSAIC: A Workload-Driven Simulation and Design-Space Exploration Framework for Heterogeneous NPUs

Arghadip Das, Hoseok Kim, Soomin Lee +3

AI model architectures are diversifying rapidly. Although dense matrix multiplication underlies today's CNNs and transformers, emerging architectures (state-space models, long conv…

cs.AR2026

BIDENT: Heterogeneous Operator-level Mapping for Efficient Edge Inference

Hoseok Kim, Arghadip Das, Soumendu Ghosh +2

Modern edge System-on-Chips (SoCs) integrate heterogeneous processing units (PUs) such as CPUs, GPUs, and NPUs, yet current inference stacks map entire models to a single PU, leavi…

cs.AR2026

SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference

Aradhana Mohan Parvathy, Soumendu Kumar Ghosh, Shamik Kundu +4

The rapid growth in sizes of Large language models (LLMs) results in high compute and memory costs during inference. Quantization has been a significant pathway to addressing this…

cs.DC2026

SlimEdge: Performance and Device Aware Distributed DNN Deployment on Resource-Constrained Edge Hardware

Mahadev Sunil Kumar, Arnab Raha, Debayan Das +3

Distributed deep neural networks (DNNs) have become central to modern computer vision, yet their deployment on resource-constrained edge devices remains hindered by substantial par…