activity
20192021
most citedA Study of BFLOAT16 for Deep Learning Training

66 citations · 107 across the 5 of their papers we have counts for

collaborators

7 papers

cs.DC2021

Extending Sparse Tensor Accelerators to Support Multiple Compression Formats

Eric Qin, Geonhwa Jeong, William Won +7

Sparsity, which occurs in both scientific applications and Deep Learning (DL) models, has been a key target of optimization within recent ASIC accelerators due to the potential mem…

cs.CL2021

The Sensitivity of Word Embeddings-based Author Detection Models to Semantic-preserving Adversarial Perturbations

Jeremiah Duncan, Fabian Fallas, Chris Gropp +11

Authorship analysis is an important subject in the field of natural language processing. It allows the detection of the most likely writer of articles, news, books, or messages. Th…

cs.DC2020

Optimizing Deep Learning Recommender Systems' Training On CPU Cluster Architectures

Dhiraj Kalamkar, Evangelos Georganas, Sudarshan Srinivasan +3

During the last two years, the goal of many researchers has been to squeeze the last bit of performance out of HPC system for AI tasks. Often this discussion is held in the context…

cs.LG2019

K-TanH: Efficient TanH For Deep Learning

Abhisek Kundu, Alex Heinecke, Dhiraj Kalamkar +7

We propose K-TanH, a novel, highly accurate, hardware efficient approximation of popular activation function TanH for Deep Learning. K-TanH consists of parameterized low-precision…

cs.DC2019

High Performance Scalable FPGA Accelerator for Deep Neural Networks

Sudarshan Srinivasan, Pradeep Janedula, Saurabh Dhoble +7

Low-precision is the first order knob for achieving higher Artificial Intelligence Operations (AI-TOPS). However the algorithmic space for sub-8-bit precision compute is diverse, w…

cs.LG201966 cited

A Study of BFLOAT16 for Deep Learning Training

Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi +16

This paper presents the first comprehensive empirical study demonstrating the efficacy of the Brain Floating Point (BFLOAT16) half-precision format for Deep Learning training acros…