13 papers
Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator
Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban
Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Judge benchmark, the top-ranked backbone alt…
ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist Deferral
Md Rabiul Islam, Samir Abdaljalil, Erchin Serpedin +1
Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imaging data and task-specific tra…
Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation
Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban
Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's language can invert backbone rankings. We loca…
IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs
Samir Abdaljalil, Erchin Serpedin, Hasan Kurban
We introduce ISOSCI, a benchmark of isomorphic cross-domain science problem pairs that separates reasoning ability from domain knowledge retrieval in LLM evaluation. Each pair shar…
4D Synchronized Fields: Motion-Language Gaussian Splatting for Temporal Scene Understanding
Mohamed Rayan Barhdadi, Samir Abdaljalil, Rasul Khanbayov +2
Current 4D representations decouple geometry, motion, and semantics: reconstruction methods discard interpretable motion structure; language-grounded methods attach semantics after…
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
Samir Abdaljalil, Erchin Serpedin, Hasan Kurban
Large language models are increasingly used to answer and verify scientific claims, yet existing evaluations typically assume that a model must always produce a definitive answer.…