activity
20192022
most citedCentaur: A Chiplet-based, Hybrid Sparse-Dense Accelerator for Personalized Recommendations

6 citations · 9 across the 4 of their papers we have counts for

collaborators

6 papers

cs.AR20221 cited

Training Personalized Recommendation Systems from (GPU) Scratch: Look Forward not Backwards

Youngeun Kwon, Minsoo Rhu

Personalized recommendation models (RecSys) are one of the most popular machine learning workload serviced by hyperscalers. A critical challenge of training RecSys is its high memo…

cs.AR2020

Tensor Casting: Co-Designing Algorithm-Architecture for Personalized Recommendation Training

Youngeun Kwon, Yunjae Lee, Minsoo Rhu

Personalized recommendations are one of the most widely deployed machine learning (ML) workload serviced from cloud datacenters. As such, architectural solutions for high-performan…

cs.DC20206 cited

Centaur: A Chiplet-based, Hybrid Sparse-Dense Accelerator for Personalized Recommendations

Ranggi Hwang, Taehun Kim, Youngeun Kwon +1

Personalized recommendations are the backbone machine learning (ML) algorithm that powers several important application domains (e.g., ads, e-commerce, etc) serviced from cloud dat…

cs.AR20191 cited

NeuMMU: Architectural Support for Efficient Address Translations in Neural Processing Units

Bongjoon Hyun, Youngeun Kwon, Yujeong Choi +2

To satisfy the compute and memory demands of deep neural networks, neural processing units (NPUs) are widely being utilized for accelerating deep learning algorithms. Similar to ho…

cs.LG2019

TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning

Youngeun Kwon, Yunjae Lee, Minsoo Rhu

Recent studies from several hyperscalars pinpoint to embedding layers as the most memory-intensive deep learning (DL) algorithm being deployed in today's datacenters. This paper ad…

cs.DC20191 cited

Beyond the Memory Wall: A Case for Memory-centric HPC System for Deep Learning

Youngeun Kwon, Minsoo Rhu

As the models and the datasets to train deep learning (DL) models scale, system architects are faced with new challenges, one of which is the memory capacity bottleneck, where the…