30 papers
Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers
Evelyn Turri, Davide Bucciarelli, Sara Sarto +2
Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape ima…
Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
Davide Lobba, Fulvio Sanguigni, Bin Ren +3
Recent advances in Virtual Try-On (VTON) and Virtual Try-Off (VTOFF) have greatly improved photo-realistic fashion synthesis and garment reconstruction. However, existing datasets…
Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing
Luca Barsellotti, Martin Sundermeyer, Mattia Segu +5
Object-centric models inspired by DETR have become the dominant paradigm for open-vocabulary video instance segmentation (OV-VIS). While recent efforts have reduced the computation…
Mind the Heads: Topological Representation Alignment for Multimodal LLMs
Davide Caffagni, Alberto Compagnoni, Federico Melis +5
Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an…
Diffusion Language Models: An Experimental Analysis
Thomas Bertolani, Davide Bucciarelli, Leonardo Zini +2
Large Language Models (LLMs) have revolutionized language modeling through autoregressive generation, enabling strong performance across a wide range of tasks. Recently, Diffusion…
Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation
Tobia Poppi, Silvia Cappelletti, Sara Sarto +5
Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interven…