7 papers
Composed Object Retrieval: Object-level Retrieval via Composed Expressions
Tong Wang, Guanyu Yang, Nian Liu +4
Retrieving fine-grained visual content based on user intent remains a challenge in multimodal systems. Although current Composed Image Retrieval (CIR) methods combine reference ima…
FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
Senmao Li, Kai Wang, Salman Khan +3
Visual Autoregressive (VAR) modeling departs from the next-token prediction paradigm of traditional Autoregressive (AR) models through next-scale prediction, enabling high-quality…
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
Ahmed Heakl, Abdelrahman M. Shaker, Youssef Mohamed +4
When a model produces a correct solution under reinforcement learning with verifiable rewards (RLVR), every token receives the same reward signal regardless of whether it was a dec…
WorldCache: Content-Aware Caching for Accelerated Video World Models
Umair Nawaz, Ahmed Heakl, Ufaq Khan +3
Diffusion Transformers (DiTs) power high-fidelity video world models but remain computationally expensive due to sequential denoising and costly spatio-temporal attention. Training…
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
Abdelrahman Shaker, Ahmed Heakl, Jaseel Muhammad +8
Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry and too heavy for deployment on…
Diversity Has Always Been There in Your Visual Autoregressive Models
Tong Wang, Guanyu Yang, Nian Liu +6
Visual Autoregressive (VAR) models have recently garnered significant attention for their innovative next-scale prediction paradigm, offering notable advantages in both inference e…