From the 1 of 7 linked papers with an AI index.
7 papers
LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving
Ming-Yen Lee, Hanchen Yang, Faaiq Waqar +4
The paper introduces LLMET, a cross‑layer simulation framework that evaluates how emerging monolithic 3D (M3D) on‑chip memory can cut energy use when serving large language models,…
Probabilistic Memory for Trustworthy Edge Intelligence
Likai Pei, Jiahao Zheng, Xueji Zhao +9
Probabilistic computation plays an important role in trustworthy edge intelligence to quantify uncertainty, enhance robustness, reconstruct data, and protect privacy, but its adopt…
ConFu: Contemplate the Future for Better Speculative Sampling
Zongyue Qin, Raghavv Goel, Mukul Gagrani +3
Speculative decoding has emerged as a powerful approach to accelerate large language model (LLM) inference by employing lightweight draft models to propose candidate tokens that ar…
NeuroSim V1.5: Improved Software Backbone for Benchmarking Compute-in-Memory Accelerators with Device and Circuit-level Non-idealities
James Read, Ming-Yen Lee, Wei-Hsing Huang +3
The exponential growth of artificial intelligence (AI) applications has exposed the inefficiency of conventional von Neumann architectures, where frequent data transfers between co…
ChatNeuroSim: An LLM Agent Framework for Automated Compute-in-Memory Accelerator Deployment and Optimization
Ming-Yen Lee, Shimeng Yu
Compute-in-Memory (CIM) architectures have been widely studied for deep neural network (DNN) acceleration by reducing data transfer overhead between the memory and computing units.…
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories
Ming-Yen Lee, Faaiq Waqar, Hanchen Yang +3
Long-context Large Language Model (LLM) inference faces increasing compute bottlenecks as attention calculations scale with context length, primarily due to the growing KV-cache tr…