6 papers
Knowledge Distillation for Visual Autoregressive Models
Elia Peruzzo, Aritra Bhowmik, Guillaume Sautiere +2
Autoregressive (AR) image generation models are highly expressive but computationally intensive, motivating effective model compression. Knowledge distillation (KD) is a natural ap…
Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference
Agata Żywot, Iason Skylitsis, Thijmen Nijdam +4
Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without r…
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
Aritra Bhowmik, Denis Korzhenkov, Cees G. M. Snoek +2
Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models…
Structured-Noise Masked Modeling for Video, Audio and Beyond
Aritra Bhowmik, Fida Mohammad Thoker, Carlos Hinojosa +2
Masked modeling has emerged as a powerful self-supervised learning framework, but existing methods largely rely on random masking, disregarding the structural properties of differe…
TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning
Aritra Bhowmik, Mohammad Mahdi Derakhshani, Dennis Koelma +3
Spatial awareness is key to enable embodied multimodal AI systems. Yet, without vast amounts of spatial supervision, current Multimodal Large Language Models (MLLMs) struggle at th…
Union-over-Intersections: Object Detection beyond Winner-Takes-All
Aritra Bhowmik, Pascal Mettes, Martin R. Oswald +1
This paper revisits the problem of predicting box locations in object detection architectures. Typically, each box proposal or box query aims to directly maximize the intersection-…