2 papers
cs.DC2026
AoiZora: Topology-Aware Auto-Parallel Optimization for Inference of Diffusion Transformers
Kaijian Wang, Yuanyuan Xu, Fanjiang Ye +5
Video diffusion has quickly grown into a key generative serving workload, yet producing each clip demands many denoising iterations over large spatio-temporal latents, which puts l…
cs.PF2026
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU
Jevin Jiang, Ying Chen, Blake A. Hechtman +2
Large Language Model (LLM) deployment is increasingly shifting to cost-efficient accelerators like Google's Tensor Processing Units (TPUs), prioritizing both performance and total…