6 papers
Binding Visual Features Point by Point
Udith Haputhanthri, Declan Campbell, Rim Assouel +2
Despite success on standard benchmarks, vision language models display persistent failures on tasks involving processing of multi-object scenes, including many tasks that are relat…
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
Rim Assouel, Amir Bar, Michal Drozdzal +1
Despite remarkable progress in Multimodal Large Language Models (MLLMs), these models still struggle with fine-grained understanding tasks. In this work, we propose Procedurally Ge…
Visual symbolic mechanisms: Emergent symbol processing in vision language models
Rim Assouel, Declan Campbell, Yoshua Bengio +1
To accurately process a visual scene, observers must bind features together to represent individual objects. This capacity is necessary, for instance, to distinguish an image conta…
The BrowserGym Ecosystem for Web Agent Research
Thibault Le Sellier De Chezelles, Maxime Gasse, Alexandre Drouin +17
The BrowserGym ecosystem addresses the growing need for efficient evaluation and benchmarking of web agents, particularly those leveraging automation and Large Language Models (LLM…
Object-centric Binding in Contrastive Language-Image Pretraining
Rim Assouel, Pietro Astolfi, Florian Bordes +2
Recent advances in vision language models (VLM) have been driven by contrastive models such as CLIP, which learn to associate visual information with their corresponding text descr…
Action abstractions for amortized sampling
Oussama Boussif, Léna Néhale Ezzine, Joseph D Viviano +6
As trajectories sampled by policies used by reinforcement learning (RL) and generative flow networks (GFlowNets) grow longer, credit assignment and exploration become more challeng…