1 paper
Kevin Yang, Violet Yao, John DeNero +1
We propose an efficient batching strategy for variable-length decoding on GPU architectures. During decoding, when candidates terminate or are pruned according to heuristics, our s…