2 papers
cs.DC2026
Enabling Memory-efficient Im2win Convolution with Multi-precision Support on GPU CUDA and Tensor Cores
Xiang Fu, Jixiang Ma, Xinpeng Zhang +3
Convolution is a principal computational bottleneck in deep neural networks, and its efficiency depends on tight integration between algorithms and GPU hardware. Existing GPU convo…
quant-ph2025
QSteed: A Resource-Virtualized and Hardware-Aware Quantum Compilation Framework for Real Quantum Computing Processors
Hong-Ze Xu, Zheng-An Wang, Yu-Long Feng +11
As quantum computing systems continue to scale up and become more clustered, efficiently compiling user quantum programs into high fidelity executable sequences on real hardware re…