most citedHalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations

2 citations · 3 across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL2026

Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator

Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban

Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Judge benchmark, the top-ranked backbone alt…

cs.CL2026

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti +11

We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question ans…

cs.CL2026

IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs

Samir Abdaljalil, Erchin Serpedin, Hasan Kurban

We introduce ISOSCI, a benchmark of isomorphic cross-domain science problem pairs that separates reasoning ability from domain knowledge retrieval in LLM evaluation. Each pair shar…

cs.CL2026

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation

Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban

Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's language can invert backbone rankings. We loca…

cs.CL20261 cited

Knowing When Not to Answer: Abstention-Aware Scientific Reasoning

Samir Abdaljalil, Erchin Serpedin, Hasan Kurban

Large language models are increasingly used to answer and verify scientific claims, yet existing evaluations typically assume that a model must always produce a definitive answer.…

cs.CL2026

Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs

Samir Abdaljalil, Parichit Sharma, Erchin Serpedin +1

Hallucinations in large language models remain a persistent challenge, particularly in multilingual and generative settings where factual consistency is difficult to maintain. Whil…