13 citations · 25 across the 5 of their papers we have counts for
5 papers
Is Flash Attention Stable?
Alicia Golden, Samuel Hsia, Fei Sun +8
Training large-scale machine learning models poses distinct system challenges, given both the size and complexity of today's workloads. Recently, many organizations training state-…
Data Acquisition: A New Frontier in Data-centric AI
Lingjiao Chen, Bilge Acun, Newsha Ardalani +8
As Machine Learning (ML) systems continue to grow, the demand for relevant and comprehensive datasets becomes imperative. There is limited study on the challenges of data acquisiti…
RecShard: Statistical Feature-Based Memory Optimization for Industry-Scale Neural Recommendation
Geet Sethi, Bilge Acun, Niket Agarwal +3
We propose RecShard, a fine-grained embedding table (EMB) partitioning and placement technique for deep learning recommendation models (DLRMs). RecShard is designed based on two ke…
TT-Rec: Tensor Train Compression for Deep Learning Recommendation Models
Chunxing Yin, Bilge Acun, Xing Liu +1
The memory capacity of embedding tables in deep learning recommendation models (DLRMs) is increasing dramatically from tens of GBs to TBs across the industry. Given the fast growth…
Understanding Training Efficiency of Deep Learning Recommendation Models at Scale
Bilge Acun, Matthew Murphy, Xiaodong Wang +3
The use of GPUs has proliferated for machine learning workflows and is now considered mainstream for many deep learning models. Meanwhile, when training state-of-the-art personal r…