1 paper · 1 filter
Xuanwen Ding, Chengjun Pan, Zejun Li +3
Evaluating multimodal large language models (MLLMs) is increasingly expensive, as the growing size and cross-modality complexity of benchmarks demand significant scoring efforts. T…