1 paper
Tim Rädsch, Leon Mayer, Simon Pavicic +8
Reliable evaluation of AI models is critical for scientific progress and practical application. While existing VLM benchmarks provide general insights into model capabilities, thei…