9 citations · 21 across the 14 of their papers we have counts for
4 papers · 1 filter
Accelerating Bandwidth-Bound Deep Learning Inference with Main-Memory Accelerators
Benjamin Y. Cho, Jeageun Jung, Mattan Erez
DL inference queries play an important role in diverse internet services and a large fraction of datacenter cycles are spent on processing DL inference queries. Specifically, the m…
WoLFRaM: Enhancing Wear-Leveling and Fault Tolerance in Resistive Memories using Programmable Address Decoders
Leonid Yavits, Lois Orosa, Suyash Mahar +4
Resistive memories have limited lifetime caused by limited write endurance and highly non-uniform write access patterns. Two main techniques to mitigate endurance-related memory fa…
Training with Multi-Layer Embeddings for Model Reduction
Benjamin Ghaemmaghami, Zihao Deng, Benjamin Cho +4
Modern recommendation systems rely on real-valued embeddings of categorical features. Increasing the dimension of embedding vectors improves model accuracy but comes at a high cost…
FlexSA: Flexible Systolic Array Architecture for Efficient Pruned DNN Model Training
Sangkug Lym, Mattan Erez
Modern deep learning models have high memory and computation cost. To make them fast and memory-cost efficient, structured model pruning is commonly used. We find that pruning a mo…