11 citations · 40 across the 8 of their papers we have counts for
4 papers · 1 filter
C-SAW: A Framework for Graph Sampling and Random Walk on GPUs
Santosh Pandey, Lingda Li, Adolfy Hoisie +2
Many applications require to learn, mine, analyze and visualize large-scale graphs. These graphs are often too large to be addressed efficiently using conventional graph processing…
EZLDA: Efficient and Scalable LDA on GPUs
Shilong Wang, Hang Liu, Anil Gaihre +1
LDA is a statistical approach for topic modeling with a wide range of applications. However, there exist very few attempts to accelerate LDA on GPUs which come with exceptional com…
FTRANS: Energy-Efficient Acceleration of Transformers using FPGA
Bingbing Li, Santosh Pandey, Haowen Fang +7
In natural language processing (NLP), the "Transformer" architecture was proposed as the first transduction model replying entirely on self-attention mechanisms without using seque…
SuperNeurons: FFT-based Gradient Sparsification in the Distributed Training of Deep Neural Networks
Linnan Wang, Wei Wu, Junyu Zhang +4
The performance and efficiency of distributed training of Deep Neural Networks highly depend on the performance of gradient averaging among all participating nodes, which is bounde…