1 paper
Taesung Kwon, Lorenzo Bianchi, Lennart Wittke +5
Recent diffusion models increasingly favor Transformer backbones, motivated by the remarkable scalability of fully attentional architectures. Yet the locality bias, parameter effic…