1 paper
Yiwei Yang, Xiangyu Gao, Yuan Zhou +3
Modern deep learning workloads often consist of many small tensor operations, especially in inference, attention, and micro-batched training. In these settings, kernel launch overh…