7 papers
On Design Principles for Efficient Heterogeneous DRAM-PIM-GPU Systems
Corey Lammie, Hadjer Benmeziane, William Andrew Simon +1
Heterogeneous DRAM-based processing-in-memory (PIM)-GPU systems promise significant efficiency gains for decode-phase large language model (LLM) inference, particularly in long-out…
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM
Zehao Fan, Garrett Gagnon, Zhenyu Liu +1
Transformer-based models are the foundation of modern machine learning, but their execution, particularly during autoregressive decoding in large language models (LLMs), places sig…
Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems
Zehao Fan, Zhenyu Liu, Yunzhen Liu +4
Mixture-of-Experts (MoE) models scale large language models through conditional computation, but inference becomes memory-bound once expert weights exceed the capacity of GPU memor…
SparseST: Exploiting Data Sparsity in Spatiotemporal Modeling and Prediction
Junfeng Wu, Hadjer Benmeziane, Kaoutar El Maghraoui +2
Spatiotemporal data mining (STDM) has a wide range of applications in various complex physical systems (CPS), i.e., transportation, manufacturing, healthcare, etc. Among all the pr…
AnalogNAS-Bench: A NAS Benchmark for Analog In-Memory Computing
Aniss Bessalah, Hatem Mohamed Abdelmoumen, Karima Benatchba +1
Analog In-memory Computing (AIMC) has emerged as a highly efficient paradigm for accelerating Deep Neural Networks (DNNs), offering significant energy and latency benefits over con…
Enhancing Downstream Analysis in Genome Sequencing: Species Classification While Basecalling
Riselda Kodra, Hadjer Benmeziane, Irem Boybat +1
The ability to quickly and accurately identify microbial species in a sample, known as metagenomic profiling, is critical across various fields, from healthcare to environmental sc…