computer architecture

CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference

arXiv:2607.13649

summary

The paper introduces CIMERA, a hardware accelerator that integrates compute-in-interconnect and memory with reconfigurable precision to run large language models more energy‑efficiently, achieving up to 25× higher energy efficiency than an Nvidia H100 for 1B‑parameter models.

Abstract

LLM impose significant computational and memory demands, creating challenges for energy-efficient inference across platforms ranging from data centers to power-constrained edge devices. Weight precision plays a critical role in balancing inference accuracy, throughput, and energy consumption, while modern LLM workloads exhibit pronounced heterogeneity and tolerance that favors adaptive precision execution. This paper presents CIMERA, a reconfigurable-precision LLM inference accelerator that integrates compute-in-interconnect and memory to mitigate the memory wall and enable precision-aware execution. Compared to Nvidia H100, CIMERA delivers up to and higher energy efficiency for 1B and 13B models, respectively.

Accepted to 2026 IEEE 8th International Conference on Artificial Intelligence Circuits and Systems (AICAS'26)

Topics & keywords

#llm inference#reconfigurable precision#compute-in-interconnect#memory‑centric accelerator#energy efficiencycompute-in-interconnectreconfigurable precisionlarge language modelenergy‑efficient inferencememory wallhardware accelerator