6 papers
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
Mayug Maniparambil, Nils Hoehing, Janak Kapuriya +5
Solving topological grid puzzles requires reasoning over global spatial invariants such as connectivity, loop closure, and region symmetry and remains challenging for even the most…
Underrepresented in Foundation Model Pretraining Data? A One-Shot Probe
Chris Vorster, Mayug Maniparambil, Noel E. O'Connor +2
Large-scale Vision-Language Foundation Models (VLFMs), such as CLIP, now underpin a wide range of computer vision research and applications. VLFMs are often adapted to various doma…
Hold-One-Shot-Out (HOSO) for Validation-Free Few-Shot CLIP Adapters
Chris Vorster, Mayug Maniparambil, Noel E. O'Connor +2
In many CLIP adaptation methods, a blending ratio hyperparameter controls the trade-off between general pretrained CLIP knowledge and the limited, dataset-specific supervision from…
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
Nils Hoehing, Mayug Maniparambil, Ellen Rushe +2
We propose RocketScience, an open-source contrastive VLM benchmark that tests for spatial relation understanding. It is comprised of entirely new real-world image-text pairs coveri…
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali +3
Recent contrastive multimodal vision-language models like CLIP have demonstrated robust open-world semantic understanding, becoming the standard image backbones for vision-language…
Pinpoint Counterfactuals: Reducing social bias in foundation models via localized counterfactual generation
Kirill Sirotkin, Marcos Escudero-Viñolo, Pablo Carballeira +3
Foundation models trained on web-scraped datasets propagate societal biases to downstream tasks. While counterfactual generation enables bias analysis, existing methods introduce a…