activity
20242026
collaborators

6 papers

cs.CV2026

Knowledge Distillation for Visual Autoregressive Models

Elia Peruzzo, Aritra Bhowmik, Guillaume Sautiere +2

Autoregressive (AR) image generation models are highly expressive but computationally intensive, motivating effective model compression. Knowledge distillation (KD) is a natural ap…

cs.CV2026

Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference

Agata Żywot, Iason Skylitsis, Thijmen Nijdam +4

Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without r…

cs.CV2025

MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models

Aritra Bhowmik, Denis Korzhenkov, Cees G. M. Snoek +2

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models…

cs.LG2025

Structured-Noise Masked Modeling for Video, Audio and Beyond

Aritra Bhowmik, Fida Mohammad Thoker, Carlos Hinojosa +2

Masked modeling has emerged as a powerful self-supervised learning framework, but existing methods largely rely on random masking, disregarding the structural properties of differe…

cs.CV2025

TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning

Aritra Bhowmik, Mohammad Mahdi Derakhshani, Dennis Koelma +3

Spatial awareness is key to enable embodied multimodal AI systems. Yet, without vast amounts of spatial supervision, current Multimodal Large Language Models (MLLMs) struggle at th…

cs.CV2024

Union-over-Intersections: Object Detection beyond Winner-Takes-All

Aritra Bhowmik, Pascal Mettes, Martin R. Oswald +1

This paper revisits the problem of predicting box locations in object detection architectures. Typically, each box proposal or box query aims to directly maximize the intersection-…