Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
ABRA: Teleporting Fine-Tuned Knowledge Across Domains for Open-Vocabulary Object Detection
Mattia Bernardi, Chiara Cappellino, Matteo Mosconi +3
Although recent Open-Vocabulary Object Detection architectures, such as Grounding DINO, demonstrate strong zero-shot capabilities, their performance degrades significantly under do…
cs.CV2025
Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
Chengxi Min, Wei Wang, Yahui Liu +4
Mixture-of-Experts (MoE) models have emerged as a promising direction for scaling vision architectures efficiently. Among them, Soft MoE improves training stability by assigning ea…
cs.CV2025
Causal Graphical Models for Vision-Language Compositional Understanding
Fiorenzo Parascandolo, Nicholas Moratelli, Enver Sangineto +2
Recent work has empirically shown that Vision-Language Models (VLMs) struggle to fully understand the compositional properties of the human language, usually modeling an image capt…