2 papers
cs.LG2025
Domain Specific Benchmarks for Evaluating Multimodal Large Language Models
Khizar Anjum, Muhammad Arbab Arshad, Kadhim Hayawi +10
Large language models (LLMs) are increasingly being deployed across disciplines due to their advanced reasoning and problem solving capabilities. To measure their effectiveness, va…
cs.AI2024
Putting GPT-4o to the Sword: A Comprehensive Evaluation of Language, Vision, Speech, and Multimodal Proficiency
Sakib Shahriar, Brady Lund, Nishith Reddy Mannuru +5
As large language models (LLMs) continue to advance, evaluating their comprehensive capabilities becomes significant for their application in various fields. This research study co…