3 papers
cs.AR2026
CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference Acceleration
Yue Jiet Chong, Yimin Wang, Zhen Wu +3
Large language model (LLM) inference exhibits substantial variability across adaptation modes, context lengths, and request concurrency, creating challenges for maintaining high ut…
cs.AR2026
CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference
Yue Jiet Chong, Yimin Wang, Wei Zhang +1
LLM impose significant computational and memory demands, creating challenges for energy-efficient inference across platforms ranging from data centers to power-constrained edge dev…
cs.ET2026
A Novel FeFET Differential Bit-Cell With Hybrid Volatile and Non-Volatile Memory Modes
Jianze Wang, Wei Zhang, Xuanyao Fong
Non-volatile SRAM (nvSRAM) designs have been investigated to address the high leakage power of CMOS-based SRAM and the large write latency of emerging non-volatile memory (eNVM) te…