1 paper
Hanqing Li, Lizhou Wu, Tiejun Li +6
Modern GPUs rely on private per-SM L1 caches and a shared L2 cache, but this organization obscures cross-SM reuse: an L1 miss is typically forwarded to L2 even when the requested l…