From the 1 of 21 linked papers with an AI index.
21 papers
AVQ-Attention: Adaptive Vector-Quantized Attention
Winfried van den dool, Patrick Forré, Amir Habibian +2
The paper introduces Adaptive Vector-Quantized (AVQ) Attention, which dynamically allocates codebook capacity to the most important regions of the key space, preserving O(MN) compl…
MobileWan: Closing the Quality Gap for Mobile Video Diffusion
Mohsen Ghafoorian, Denis Korzhenkov, Adil Karjauv +9
Recent advances in video diffusion have been driven by scaling transformer-based architectures to billions of parameters, substantially improving visual fidelity and motion coheren…
Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation
Long Vu, Tan Ngo, Animesh Karnewar +5
Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and…
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
Yanlong Chen, Amirhossein Habibian, Luca Benini +1
Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes significant accuracy loss. Despite its pot…
Knowledge Distillation for Visual Autoregressive Models
Elia Peruzzo, Aritra Bhowmik, Guillaume Sautiere +2
Autoregressive (AR) image generation models are highly expressive but computationally intensive, motivating effective model compression. Knowledge distillation (KD) is a natural ap…
Multi-Scale Local Speculative Decoding for Image Generation
Elia Peruzzo, Guillaume Sautière, Amirhossein Habibian
Autoregressive (AR) models have achieved remarkable success in image synthesis, yet their sequential nature imposes significant latency constraints. Speculative Decoding offers a p…