gpu inference performance 1kernel optimization 1nvlink vs pcie 1quantization 1runtime analysis 1sharding 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
The Label Complexity of Class-Conditional Coverage under Distribution Shift
Weijia Han, Lisha Qu
Conformal prediction certifies that a classifier's prediction sets cover the truth, and that certificate is marginal. Many recognition benchmarks build distribution shift into eval…
cs.DC2026
Decomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUs
Weijia Han, Lisha Qu
The paper analyzes how much of the reported speedup from quantized inference on NVIDIA RTX A5000 GPUs comes from reduced runtime versus kernel and quantization changes, using a mat…