collaborators

6 papers

cs.AR2026

LLM Inference on IMC-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism

Yimin Wang, Yue Jiet Chong, Xuanyao Fong

LLM inference has become an essential service, yet it imposes unprecedented demands on memory bandwidth, computational density, and communication efficiency. While IMC is a promisi…

cs.AR2026

CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference Acceleration

Yue Jiet Chong, Yimin Wang, Zhen Wu +3

Large language model (LLM) inference exhibits substantial variability across adaptation modes, context lengths, and request concurrency, creating challenges for maintaining high ut…

cs.AR2026

CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference

Yue Jiet Chong, Yimin Wang, Wei Zhang +1

LLM impose significant computational and memory demands, creating challenges for energy-efficient inference across platforms ranging from data centers to power-constrained edge dev…

cs.AR2026

PRIMAL: Processing-In-Memory Based Low-Rank Adaptation for LLM Inference Accelerator

Yue Jiet Chong, Yimin Wang, Zhen Wu +1

This paper presents PRIMAL, a processing-in-memory (PIM) based large language model (LLM) inference accelerator with low-rank adaptation (LoRA). PRIMAL integrates heterogeneous PIM…

cs.AR2025

PICNIC: Silicon Photonic Interconnected Chiplets with Computational Network and In-memory Computing for LLM Inference Acceleration

Yue Jiet Chong, Yimin Wang, Zhen Wu +1

This paper presents a 3D-stacked chiplets based large language model (LLM) inference accelerator, consisting of non-volatile in-memory-computing processing elements (PEs) and Inter…

cs.AR2025

LEAP: LLM Inference on Scalable PIM-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism

Yimin Wang, Yue Jiet Chong, Xuanyao Fong

Large language model (LLM) inference has been a prevalent demand in daily life and industries. The large tensor sizes and computing complexities in LLMs have brought challenges to…