Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
Anatole Jacquin de Margerie, Alexis Roger, Irina Rish
Reproducibility remains a cornerstone of scientific progress, yet complex multimodal models often lack transparent implementation details and accessible training infrastructure. In…
cs.CV2025
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models
Alexis Roger, Prateek Humane, Daniel Z. Kaplan +7
The proliferation of Vision-Language Models (VLMs) in the past several years calls for rigorous and comprehensive evaluation methods and benchmarks. This work analyzes existing VLM…
cs.CV2022
Aligning MAGMA by Few-Shot Learning and Finetuning
Jean-Charles Layoun, Alexis Roger, Irina Rish
The goal of vision-language modeling is to allow models to tie language understanding with visual inputs. The aim of this paper is to evaluate and align the Visual Language Model (…