1 paper
ChiHeng Jin, Hongche Yu, Xihui Chen
Large language model (LLM) decoding is latency-sensitive and often bottlenecked by fragmented operator execution and repeated off-chip materialization of intermediate tensors. Prio…