◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Ganesh Bikshandi

3 papers hereh-index 111k citations16 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author1
  • first author1
  • middle author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • cs.DC1

identity via Semantic Scholar / OpenAlex

activity
20232026
most citedFlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

16 citations · 16 across the 2 of their papers we have counts for

collaborators

3 papers

cs.DC2026

Hardware-Aware Reformulation of Convolutions for Efficient Execution on Specialized AI Hardware: A Case Study on NVIDIA Tensor Cores

Ganesh Bikshandi

Convolutional Neural Networks (CNNs) are central to modern AI, but their performance is often limited by hardware constraints. NVIDIA Tensor Cores, for instance, require input chan…

cs.LG2024★ 16 cited

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Jay Shah, Ganesh Bikshandi, Ying Zhang +3

Attention, as a core layer of the ubiquitous Transformer architecture, is the bottleneck for large language models and long-context applications. FlashAttention elaborated an appro…

cs.LG2023

A Case Study in CUDA Kernel Fusion: Implementing FlashAttention-2 on NVIDIA Hopper Architecture using the CUTLASS Library

Ganesh Bikshandi, Jay Shah

We provide an optimized implementation of the forward pass of FlashAttention-2, a popular memory-aware scaled dot-product attention algorithm, as a custom fused CUDA kernel targeti…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.