1 paper · 1 filter
Mengting He, Shihao Xia, Haomin Jia +2
Large language model (LLM) inference systems rely on CUDA kernels for core GPU computations, yet the interface between models and kernels is implicit and poorly specified. Models a…