works on

From the 1 of 5 linked papers with an AI index.

activity
20242026
collaborators

5 papers

cs.AR2026

CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference

Yue Jiet Chong, Yimin Wang, Wei Zhang +1

The paper introduces CIMERA, a hardware accelerator that integrates compute-in-interconnect and memory with reconfigurable precision to run large language models more energy‑effici…

cs.AR2026

PRIMAL: Processing-In-Memory Based Low-Rank Adaptation for LLM Inference Accelerator

Yue Jiet Chong, Yimin Wang, Zhen Wu +1

This paper presents PRIMAL, a processing-in-memory (PIM) based large language model (LLM) inference accelerator with low-rank adaptation (LoRA). PRIMAL integrates heterogeneous PIM…

cs.AR2025

PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration

Yue Jiet Chong, Yimin Wang, Zhen Wu +1

This paper presents a 3D-stacked chiplets based large language model (LLM) inference accelerator, consisting of non-volatile in-memory-computing processing elements (PEs) and Inter…

cs.AR2025

LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism

Yimin Wang, Yue Jiet Chong, Xuanyao Fong

Large language model (LLM) inference has been a prevalent demand in daily life and industries. The large tensor sizes and computing complexities in LLMs have brought challenges to…

cs.AI2024

Analysis of Higher-Order Ising Hamiltonians

Yunuo Cen, Zhiwei Zhang, Zixuan Wang +2

It is challenging to scale Ising machines for industrial-level problems due to algorithm or hardware limitations. Although higher-order Ising models provide a more compact encoding…