3 papers
cs.CV2026
Auto-Comp: An Automated Pipeline for Scalable Compositional Probing of Contrastive Vision-Language Models
Cristian Sbrolli, Matteo Matteucci, Toshihiko Yamasaki
Modern Vision-Language Models (VLMs) exhibit a critical flaw in compositional reasoning, often confusing "a red cube and a blue sphere" with "a blue cube and a red sphere". Disenta…
cs.CV2026
PolyGen: Fully Synthetic Vision-Language Training via Multi-Generator Ensembles
Leonardo Brusini, Cristian Sbrolli, Eugenio Lomurno +2
Synthetic data offers a scalable solution for vision-language pre-training, yet current state-of-the-art methods typically rely on scaling up a single generative backbone, which in…
cs.LG2025
From Offline to Online Memory-Free and Task-Free Continual Learning via Fine-Grained Hypergradients
Nicolas Michel, Maorong Wang, Jiangpeng He +1
Continual Learning (CL) aims to learn from a non-stationary data stream where the underlying distribution changes over time. While recent advances have produced efficient memory-fr…