Showing cs.ARShow all
3 papers · 1 filter
cs.AR2026
BIDENT: Heterogeneous Operator-level Mapping for Efficient Edge Inference
Hoseok Kim, Arghadip Das, Soumendu Ghosh +2
Modern edge System-on-Chips (SoCs) integrate heterogeneous processing units (PUs) such as CPUs, GPUs, and NPUs, yet current inference stacks map entire models to a single PU, leavi…
cs.AR2026
SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference
Aradhana Mohan Parvathy, Soumendu Kumar Ghosh, Shamik Kundu +4
The rapid growth in sizes of Large language models (LLMs) results in high compute and memory costs during inference. Quantization has been a significant pathway to addressing this…
cs.AR2024
FlexNN: A Dataflow-aware Flexible Deep Learning Accelerator for Energy-Efficient Edge Devices
Arnab Raha, Deepak A. Mathaikutty, Soumendu K. Ghosh +1
This paper introduces FlexNN, a Flexible Neural Network accelerator, which adopts agile design principles to enable versatile dataflows, enhancing energy efficiency. Unlike convent…