21 citations · 33 across the 5 of their papers we have counts for
Showing cs.PFShow all
2 papers · 1 filter
cs.PF2026
KernelSight-LM: A Kernel-Level LLM Inference Simulator
Xiteng Yao, Taeho Kim, Hengzhi Pei +7
As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to m…
cs.PF2023★ 12 cited
Tensor Slicing and Optimization for Multicore NPUs
Rafael Sousa, Marcio Pereira, Yongin Kwon +5
Although code generation for Convolution Neural Network (CNN) models has been extensively studied, performing efficient data slicing and parallelization for highly-constrai\-ned Mu…