1 citations · 2 across the 3 of their papers we have counts for
3 papers
HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
Krish Agarwal, Rishi Astra, Adnan Hoque +4
We present HadaCore, a modified Fast Walsh-Hadamard Transform (FWHT) algorithm optimized for the Tensor Cores present in modern GPU hardware. HadaCore follows the recursive structu…
Accelerating a Triton Fused Kernel for W4A16 Quantized Inference with SplitK work decomposition
Adnan Hoque, Less Wright, Chih-Chieh Yang +2
We propose an implementation of an efficient fused matrix multiplication kernel for W4A16 quantized inference, where we perform dequantization and GEMM in a fused kernel using a Sp…
TP-Aware Dequantization
Adnan Hoque, Mudhakar Srivatsa, Chih-Chieh Yang +1
In this paper, we present a novel method that reduces model inference latency during distributed deployment of Large Language Models (LLMs). Our contribution is an optimized infere…