most citedTraining with Multi-Layer Embeddings for Model Reduction

5 citations · 10 across the 4 of their papers we have counts for

collaborators

6 papers

cs.AR2020

WoLFRaM: Enhancing Wear-Leveling and Fault Tolerance in Resistive Memories using Programmable Address Decoders

Leonid Yavits, Lois Orosa, Suyash Mahar +4

Resistive memories have limited lifetime caused by limited write endurance and highly non-uniform write access patterns. Two main techniques to mitigate endurance-related memory fa…

cs.LG20205 cited

Training with Multi-Layer Embeddings for Model Reduction

Benjamin Ghaemmaghami, Zihao Deng, Benjamin Cho +4

Modern recommendation systems rely on real-valued embeddings of categorical features. Increasing the dimension of embedding vectors improves model accuracy but comes at a high cost…

cs.LG20205 cited

FlexSA: Flexible Systolic Array Architecture for Efficient Pruned DNN Model Training

Sangkug Lym, Mattan Erez

Modern deep learning models have high memory and computation cost. To make them fast and memory-cost efficient, structured model pruning is commonly used. We find that pruning a mo…

cs.AR2019

Buddy Compression: Enabling Larger Memory for Deep Learning and HPC Workloads on GPUs

Esha Choukse, Michael Sullivan, Mike O'Connor +4

GPUs offer orders-of-magnitude higher memory bandwidth than traditional CPU-only systems. However, GPU device memory tends to be relatively small and the memory capacity can not be…

cs.DC2019

DeLTA: GPU Performance Model for Deep Learning Applications with In-depth Memory System Traffic Analysis

Sangkug Lym, Donghyuk Lee, Mike O'Connor +2

Training convolutional neural networks (CNNs) requires intense compute throughput and high memory bandwidth. Especially, convolution layers account for the majority of the executio…

cs.LG2019

PruneTrain: Fast Neural Network Training by Dynamic Sparse Model Reconfiguration

Sangkug Lym, Esha Choukse, Siavash Zangeneh +3

State-of-the-art convolutional neural networks (CNNs) used in vision applications have large models with numerous weights. Training these models is very compute- and memory-resourc…