From the 1 of 4 linked papers with an AI index.
4 papers
NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement
Sookyung Choi, Seungyong Lee, Kangkyu Park +11
The paper introduces NELSSA, a system that combines GPUs with processing‑near‑memory (PNM) accelerators to efficiently serve large language model requests of varying context length…
AXLE: Coordinated Offloading with Asynchronous Back-Streaming in Computational Memory Systems
Suyeon Lee, Kangkyu Park, Kwangsik Shin +1
CXL-based Computational Memory (CCM) enables near-memory processing within expanded remote memory, offering opportunities to address data movement costs in disaggregated memory sys…
Cocoon: A System Architecture for Differentially Private Training with Correlated Noises
Donghwan Kim, Xin Gu, Jinho Baek +6
Machine learning (ML) models memorize and leak training data, causing serious privacy issues to data owners. Training algorithms with differential privacy (DP), such as DP-SGD, hav…
Tuning Fast Memory Size based on Modeling of Page Migration for Tiered Memory
Shangye Chen, Jin Huang, Shuangyan Yang +7
Tiered memory, built upon a combination of fast memory and slow memory, provides a cost-effective solution to meet ever-increasing requirements from emerging applications for large…