MAST: Mask-Guided Attention Control for Training-Free Regional-Multi Style Transfer
arXiv:2604.12281
Abstract
Style transfer applies the appearance of a reference image to a content image while preserving its spatial structure. Recent diffusion-based methods achieve strong stylization but typically assume a single global style. We instead consider regional-multi style transfer, which assigns multiple references to user-specified regions of a content image. Extending them to this setting reveals two coupled shared-attention issues: ambiguous mass allocation among content and style partitions and degraded selectivity as more styles are jointly normalized, while style aggregation at the attention output further suppresses fine details. We propose MAST (Mask-Guided Attention Control for Training-Free Regional-Multi Style Transfer), a unified attention-control framework for frozen diffusion models. Logit-level Attention Mass Allocation enforces mask-derived partition masses, Sharpness-aware Temperature Scaling adaptively restores selectivity, and Discrepancy-aware Detail Injection recovers high-frequency content. MAST jointly processes all style--mask pairs in a single denoising pass without training, optimization, or post-hoc composition. Across two to five styles, MAST achieves the best average ArtFID, FID, and R-FID among all baselines, demonstrating regional style fidelity, content preservation, and scalability.
23 pages, 17 figures, 13 tables