paper

HADIS: A Hybrid Architecture for Query-Aware Diffusion Model Serving

arXiv:2509.00642

Abstract

Query-aware model serving routes queries through cascades of increasing cost to balance quality and throughput. Existing cascade systems force all queries through the lightweight stage, wasting resources on hard queries, and fix the model pair regardless of workload dynamics. In text-to-image diffusion serving, this waste is severe: quality is assessable only after full generation completes, and rejected lightweight outputs cannot be reused by the heavyweight model. HADIS is a diffusion model serving system built on a hybrid cascade architecture that combines a pre-generation router (bypassing the lightweight model for predicted-hard queries) with a post-generation discriminator (catching false-easy cases). The hybrid architecture creates a coupled optimization space: routing thresholds, model-pair selection, and GPU allocation must be jointly adapted as workloads shift. HADIS reduces this space by establishing that two-model cascades match the Pareto frontier of deeper cascades, pruning candidates into a compact Pareto-optimal configuration table offline, and re-optimizing allocation at runtime via an online MILP. The hybrid design dominates both router-only and discriminator-only alternatives under mild accuracy conditions. On a 16-GPU cluster with real-world traces, HADIS improves response quality by up to 35% and reduces SLO violations by 2.7-45 compared to baselines.

15 pages, 14 figures