6 citations · 14 across the 5 of their papers we have counts for
8 papers
A Survey of Multi-Tenant Deep Learning Inference on GPU
Fuxun Yu, Di Wang, Longfei Shangguan +3
Deep Learning (DL) models have achieved superior performance. Meanwhile, computing hardware like NVIDIA GPUs also demonstrated strong computing scaling trends with 2x throughput an…
Maximizing Communication Efficiency for Large-scale Training via 0/1 Adam
Yucheng Lu, Conglong Li, Minjia Zhang +2
1-bit gradient compression and local steps are two representative techniques that enable drastic communication reduction in distributed SGD. Their benefits, however, remain an open…
Speed-ANN: Low-Latency and High-Accuracy Nearest Neighbor Search via Intra-Query Parallelism
Zhen Peng, Minjia Zhang, Kai Li +2
Nearest Neighbor Search (NNS) has recently drawn a rapid increase of interest due to its core role in managing high-dimensional vector data in data science and AI applications. The…
ScaLA: Accelerating Adaptation of Pre-Trained Transformer-Based Language Models via Efficient Large-Batch Adversarial Noise
Minjia Zhang, Niranjan Uma Naresh, Yuxiong He
In recent years, large pre-trained Transformer-based language models have led to dramatic improvements in many natural language understanding tasks. To train these models with incr…
NxMTransformer: Semi-Structured Sparsification for Natural Language Understanding via ADMM
Connor Holmes, Minjia Zhang, Yuxiong He +1
Natural Language Processing (NLP) has recently achieved success by using huge pre-trained Transformer networks. However, these models often contain hundreds of millions or even bil…
Understanding and Generalizing Monotonic Proximity Graphs for Approximate Nearest Neighbor Search
Dantong Zhu, Minjia Zhang
Graph-based algorithms have shown great empirical potential for the approximate nearest neighbor (ANN) search problem. Currently, graph-based ANN search algorithms are designed mai…