1 paper · 1 filter
Mohammad Siavashi, Gerald Q. Maguire, Dejan Kostic +1
A single SM utilization percentage can make an LLM inference workload look compute-saturated while hiding how much useful work is being done. The problem is not that the counter is…