activity
20172024
most citedSCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks

125 citations · 136 across the 8 of their papers we have counts for

collaborators

14 papers

cs.AR2024

PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems

Dongjae Lee, Bongjoon Hyun, Taehun Kim +1

Processing-in-memory (PIM) has emerged as a promising solution for accelerating memory-intensive workloads as they provide high memory bandwidth to the processing units. This appro…

cs.AR20222 cited

SmartSAGE: Training Large-scale Graph Neural Networks using In-Storage Processing Architectures

Yunjae Lee, Jinha Chung, Minsoo Rhu

Graph neural networks (GNNs) can extract features by learning both the representation of each objects (i.e., graph nodes) and the relationship across different objects (i.e., the e…

cs.AR20221 cited

Training Personalized Recommendation Systems from (GPU) Scratch: Look Forward not Backwards

Youngeun Kwon, Minsoo Rhu

Personalized recommendation models (RecSys) are one of the most popular machine learning workload serviced by hyperscalers. A critical challenge of training RecSys is its high memo…

cs.DC2022

PARIS and ELSA: An Elastic Scheduling Algorithm for Reconfigurable Multi-GPU Inference Servers

Yunseong Kim, Yujeong Choi, Minsoo Rhu

In cloud machine learning (ML) inference systems, providing low latency to end-users is of utmost importance. However, maximizing server utilization and system throughput is also c…

cs.DC2020

LazyBatching: An SLA-aware Batching System for Cloud Machine Learning Inference

Yujeong Choi, Yunseong Kim, Minsoo Rhu

In cloud ML inference systems, batching is an essential technique to increase throughput which helps optimize total-cost-of-ownership. Prior graph batching combines the individual…

cs.AR2020

Tensor Casting: Co-Designing Algorithm-Architecture for Personalized Recommendation Training

Youngeun Kwon, Yunjae Lee, Minsoo Rhu

Personalized recommendations are one of the most widely deployed machine learning (ML) workload serviced from cloud datacenters. As such, architectural solutions for high-performan…