3 papers
cs.CV2026
Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference
Agata Żywot, Iason Skylitsis, Thijmen Nijdam +4
Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without r…
cs.CV2026
Do All Vision Transformers Need Registers? A Cross-Architectural Reassessment
Spiros Baxevanakis, Platon Karageorgis, Ioannis Dravilas +1
Training Vision Transformers (ViTs) presents significant challenges, one of which is the emergence of artifacts in attention maps, hindering their interpretability. Darcet et al. (…
cs.SD2025
Linear RNNs for autoregressive generation of long music samples
Konrad Szewczyk, Daniel Gallo Fernández, James Townsend
Directly learning to generate audio waveforms in an autoregressive manner is a challenging task, due to the length of the raw sequences and the existence of important structure on…