From the 1 of 5 linked papers with an AI index.
5 papers
CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference
Yue Jiet Chong, Yimin Wang, Wei Zhang +1
The paper introduces CIMERA, a hardware accelerator that integrates compute-in-interconnect and memory with reconfigurable precision to run large language models more energy‑effici…
PRIMAL: Processing-In-Memory Based Low-Rank Adaptation for LLM Inference Accelerator
Yue Jiet Chong, Yimin Wang, Zhen Wu +1
This paper presents PRIMAL, a processing-in-memory (PIM) based large language model (LLM) inference accelerator with low-rank adaptation (LoRA). PRIMAL integrates heterogeneous PIM…
PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration
Yue Jiet Chong, Yimin Wang, Zhen Wu +1
This paper presents a 3D-stacked chiplets based large language model (LLM) inference accelerator, consisting of non-volatile in-memory-computing processing elements (PEs) and Inter…
LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
Yimin Wang, Yue Jiet Chong, Xuanyao Fong
Large language model (LLM) inference has been a prevalent demand in daily life and industries. The large tensor sizes and computing complexities in LLMs have brought challenges to…
Analysis of Higher-Order Ising Hamiltonians
Yunuo Cen, Zhiwei Zhang, Zixuan Wang +2
It is challenging to scale Ising machines for industrial-level problems due to algorithm or hardware limitations. Although higher-order Ising models provide a more compact encoding…