3 papers
cs.PF2026
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
Yifan Zhu, Yekai Pan, Chen Ding
High-performance attention kernels are essential for Large Language Models. This paper presents analysis of CuTile-based Flash Attention memory behavior and a technique to improve…
cs.AR2025
SemanticBBV: A Semantic Signature for Cross-Program Knowledge Reuse in Microarchitecture Simulation
Zhenguo Liu, Chengao Shi, Chen Ding +1
For decades, sampling-based techniques have been the de facto standard for accelerating microarchitecture simulation, with the Basic Block Vector (BBV) serving as the cornerstone p…
cs.LG2023
Scalable CP Decomposition for Tensor Learning using GPU Tensor Cores
Zeliang Zhang, Zhuo Liu, Susan Liang +4
CP decomposition is a powerful tool for data science, especially gene analysis, deep learning, and quantum computation. However, the application of tensor decomposition is largely…