2 papers
cs.AR2024
Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
Haibin Wu, Wenming Li, Kai Yan +9
Recent neural networks (NNs) with self-attention exhibit competitiveness across different AI domains, but the essential attention mechanism brings massive computation and memory de…
cs.AR2024
Accelerating Mini-batch HGNN Training by Reducing CUDA Kernels
Meng Wu, Jingkai Qiu, Mingyu Yan +5
Heterogeneous graph neural networks (HGNNs) are essential for capturing the structure and semantic information in heterogeneous graphs. However, existing GPU-based solutions, such…