most citedAccelerating Deep Learning Model Inference on Arm CPUs with Ultra-Low Bit Quantization and Runtime

3 citations · 5 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20231 cited

Accelerating Deep Neural Networks via Semi-Structured Activation Sparsity

Matteo Grimaldi, Darshan C. Ganji, Ivan Lazarevich +1

The demand for efficient processing of deep neural networks (DNNs) on embedded devices is a significant challenge limiting their deployment. Exploiting sparsity in the network's fe…

cs.LG2023

DeepliteRT: Computer Vision at the Edge

Saad Ashfaq, Alexander Hoffman, Saptarshi Mitra +3

The proliferation of edge devices has unlocked unprecedented opportunities for deep learning model deployment in computer vision applications. However, these complex models require…

cs.CV20231 cited

YOLOBench: Benchmarking Efficient Object Detectors on Embedded Systems

Ivan Lazarevich, Matteo Grimaldi, Ravish Kumar +3

We present YOLOBench, a benchmark comprised of 550+ YOLO-based object detection models on 4 different datasets and 4 different embedded hardware platforms (x86 CPU, ARM CPU, Nvidia…

cs.LG2023

DeepGEMM: Accelerated Ultra Low-Precision Inference on CPU Architectures using Lookup Tables

Darshan C. Ganji, Saad Ashfaq, Ehsan Saboori +6

A lot of recent progress has been made in ultra low-bit quantization, promising significant improvements in latency, memory footprint and energy consumption on edge devices. Quanti…

cs.LG20223 cited

Accelerating Deep Learning Model Inference on Arm CPUs with Ultra-Low Bit Quantization and Runtime

Saad Ashfaq, MohammadHossein AskariHemmat, Sudhakar Sah +3

Deep Learning has been one of the most disruptive technological advancements in recent times. The high performance of deep learning models comes at the expense of high computationa…