papers
Publications (2)
cs.AR2025
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories
Ming-Yen Lee, Faaiq Waqar, Hanchen Yang +3
Long-context Large Language Model (LLM) inference faces increasing compute bottlenecks as attention calculations scale with context length, primarily due to the growing KV-cache tr…
cs.AR2026
LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving
Ming-Yen Lee, Hanchen Yang, Faaiq Waqar +4
The paper introduces LLMET, a cross‑layer simulation framework that evaluates how emerging monolithic 3D (M3D) on‑chip memory can cut energy use when serving large language models,…
#large language models#energy efficiency#emerging memory#monolithic 3d integration