From the 1 of 3 linked papers with an AI index.
1 paper · 1 filter
Harsh Joshi, Gautam Siddharth Kashyap, Rafiq Ali +5
The paper introduces ARGUS-EVAL, a framework that assesses vision-language models on both capability and reliability across domains, and uses it to compare several VLMs on retrieva…