1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.DC2026
Enabling Memory-efficient Im2win Convolution with Multi-precision Support on GPU CUDA and Tensor Cores
Xiang Fu, Jixiang Ma, Xinpeng Zhang +3
Convolution is a principal computational bottleneck in deep neural networks, and its efficiency depends on tight integration between algorithms and GPU hardware. Existing GPU convo…
cs.LG2024★ 1 cited
High Performance Im2win and Direct Convolutions using Three Tensor Layouts on SIMD Architectures
Xiang Fu, Xinpeng Zhang, Jixiang Ma +3
Convolution is the core component within deep neural networks and it is computationally intensive and time consuming. Tensor data layouts significantly impact convolution operation…