2 papers
cs.LG2026
TetriServe: Efficiently Serving Mixed DiT Workloads
Runyu Lu, Shiqi He, Wenxuan Tan +5
Diffusion Transformer (DiT) models excel at generating high-quality images through iterative denoising steps, but serving them under strict Service Level Objectives (SLOs) is chall…
cs.LG2026
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
Peiyuan Zhang, Matthew Noto, Wenxuan Tan +4
Achieving reliable 4-bit attention is a prerequisite for end-to-end FP4 computation on emerging FP4-capable GPUs, yet attention remains the main obstacle due to FP4's tiny dynamic…