5 papers
A Runtime-Adaptive Transformer Neural Network Accelerator on FPGAs
Ehsan Kabir, Jason D. Bakos, David Andrews +1
Transformer neural networks (TNN) excel in natural language processing (NLP), machine translation, and computer vision (CV) without relying on recurrent or convolutional layers. Ho…
FAMOUS: Flexible Accelerator for the Attention Mechanism of Transformer on UltraScale+ FPGAs
Ehsan Kabir, Md. Arafat Kabir, Austin R. J. Downey +3
Transformer neural networks (TNNs) are being applied across a widening range of application domains, including natural language processing (NLP), machine translation, and computer…
N-TORC: Native Tensor Optimizer for Real-time Constraints
Suyash Vardhan Singh, Iftakhar Ahmad, David Andrews +3
Compared to overlay-based tensor architectures like VTA or Gemmini, compilers that directly translate machine learning models into a dataflow architecture as HLS code, such as HLS4…
The BRAM is the Limit: Shattering Myths, Shaping Standards, and Building Scalable PIM Accelerators
MD Arafat Kabir, Tendayi Kamucheka, Nathaniel Fredricks +4
Many recent FPGA-based Processor-in-Memory (PIM) architectures have appeared with promises of impressive levels of parallelism but with performance that falls short of expectations…
IMAGine: An In-Memory Accelerated GEMV Engine Overlay
MD Arafat Kabir, Tendayi Kamucheka, Nathaniel Fredricks +4
Processor-in-Memory (PIM) overlays and new redesigned reconfigurable tile fabrics have been proposed to eliminate the von Neumann bottleneck and enable processing performance to sc…