Showing cs.ARShow all
2 papers · 1 filter
cs.AR2026
CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference Acceleration
Yue Jiet Chong, Yimin Wang, Zhen Wu +3
Large language model (LLM) inference exhibits substantial variability across adaptation modes, context lengths, and request concurrency, creating challenges for maintaining high ut…
cs.AR2026
CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference
Yue Jiet Chong, Yimin Wang, Wei Zhang +1
LLM impose significant computational and memory demands, creating challenges for energy-efficient inference across platforms ranging from data centers to power-constrained edge dev…