works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.AR2026

LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving

Ming-Yen Lee, Hanchen Yang, Faaiq Waqar +4

The paper introduces LLMET, a cross‑layer simulation framework that evaluates how emerging monolithic 3D (M3D) on‑chip memory can cut energy use when serving large language models,…

cs.AR2025

Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale Systems

Wei-Hsing Huang, Jianwei Jia, Yuyao Kong +4

Recent developments have introduced Kolmogorov-Arnold Networks (KAN), an innovative architectural paradigm capable of replicating conventional deep neural network (DNN) capabilitie…

cs.AR2025

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories

Ming-Yen Lee, Faaiq Waqar, Hanchen Yang +3

Long-context Large Language Model (LLM) inference faces increasing compute bottlenecks as attention calculations scale with context length, primarily due to the growing KV-cache tr…

cs.AR2025

A3D-MoE: Acceleration of Large Language Models with Mixture of Experts via 3D Heterogeneous Integration

Wei-Hsing Huang, Janak Sharda, Cheng-Jhih Shih +6

Conventional large language models (LLMs) are equipped with dozens of GB to TB of model parameters, making inference highly energy-intensive and costly as all the weights need to b…

cs.ET2025

CMOS+X: Stacking Persistent Embedded Memories based on Oxide Transistors upon GPGPU Platforms

Faaiq Waqar, Ming-Yen Lee, Seongwon Yoon +2

In contemporary general-purpose graphics processing units (GPGPUs), the continued increase in raw arithmetic throughput is constrained by the capabilities of the register file (sin…

cs.ET2025

Optimization and Benchmarking of Monolithically Stackable Gain Cell Memory for Last-Level Cache

Faaiq Waqar, Jungyoun Kwak, Junmo Lee +4

The Last Level Cache (LLC) is the processor's critical bridge between on-chip and off-chip memory levels - optimized for high density, high bandwidth, and low operation energy. To…