2 papers
cs.CL2026
EduArt: An educational-level benchmark for evaluating art history knowledge in large language models
Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza
Large language models now score near ceiling on general benchmarks, but these aggregate measures reveal little about how models behave within single disciplines. Existing art-focus…
cs.CV2025
Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography
Gianmarco Spinaci, Lukas Klic, Giovanni Colavizza
This study evaluates the capabilities of Multimodal Large Language Models (LLMs) and Vision Language Models (VLMs) in the task of single-label classification of Christian Iconograp…