5 papers
Benchmarking Frontier Text-to-Image Models on Image-Description Prompts
Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmed Rashad
Text-to-image models are typically reported on average-case prompts, which understates the gap between systems on compositionally demanding requests involving precise object counts…
Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification
Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmad ElShiekh +1
We benchmark eleven audio classification methods: five task-aware closed-set LLMs (four Gemini models plus open-weight Kimi-Audio-7B-Instruct), four fixed-vocabulary taggers (YAMNe…
Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth
Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmad ElShiekh +2
The cost of human expert evaluation is a principal bottleneck to deploying language models in specialized, high-stakes domains. This is particularly acute for Arabic sociolinguisti…
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
Sajjad Abdoli, Ghassan Al-Sumaidaee, Clayton W. Taylor +2
Code-switching -- the natural alternation between two languages within a single utterance -- remains one of the most challenging and under-studied conditions for automatic speech r…
Understanding AI Evaluation Patterns: How Different GPT Models Assess Vision-Language Descriptions
Sajjad Abdoli, Rudi Cilibrasi, Rima Al-Shikh
As AI systems increasingly evaluate other AI outputs, understanding their assessment behavior becomes crucial for preventing cascading biases. This study analyzes vision-language d…