1 paper
Swapnil Gandhi, Siva Hari, William J. Dally +1
Recent methods expose intra-request parallelism in LLM outputs, allowing independent branches to decode concurrently. Existing serving systems execute these branches eagerly or und…