gpu simulation 1hardware-software co-design 1llm inference 1performance modeling 1tile-centric modeling 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.AR2026
A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation
Zihan Liu, Jingwen Leng, Yangjie Zhou +12
Modern GPUs increasingly integrate Tensor Cores into the execution pipeline. Although aggregate tensor throughput continues to grow, aided by an operand supply that has evolved fro…
cs.DC2026
GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-Design
Yitong Ding, Jiawei Huang, Renyang Guan +7
The paper introduces GPU‑Tile‑Sim, a tile‑centric GPU simulation framework that models LLM kernels using warp‑level tile graphs to capture dependency and overlap, achieving low err…
cs.AR2026
Cache-Resident LLM Inference in GB-Scale Last-Level Caches
Wanning Zhang, Tongzhou Gu, Marco Canini +2
Large language model (LLM) inference is increasingly dominated by data movement across the memory hierarchy. Recent 3D-stacked cache technologies have enabled GB-scale last-level c…