activity
20242026
most citedHintEval: An Open-Source Python Toolkit for Hint Generation and Hint Evaluation

3 citations · 7 across the 26 of their papers we have counts for

collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL2026

Large Language Models Systematically Favor Popular Options: Evidence and Mitigation Across MCQs

Abdelrahman Abdallah, Mohammed Ali, Bhawna Piryani +2

Multiple-choice questions (MCQs) are a standard format for evaluating large language models (LLMs), yet the popularity of answer options can confound evaluation. Modern LLMs system…

cs.CL2025

DeAR: Dual-Stage Document Reranking with Reasoning Agents via LLM Distillation

Abdelrahman Abdallah, Jamshid Mozafari, Bhawna Piryani +1

Large Language Models (LLMs) have transformed listwise document reranking by enabling global reasoning over candidate sets, yet single models often struggle to balance fine-grained…

cs.CL2025

How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models

Abdelrahman Abdallah, Bhawna Piryani, Jamshid Mozafari +2

In this work, we present a systematic and comprehensive empirical evaluation of state-of-the-art reranking methods, encompassing large language model (LLM)-based, lightweight conte…

cs.CL2025

It's High Time: A Survey of Temporal Question Answering

Bhawna Piryani, Abdelrahman Abdallah, Jamshid Mozafari +2

Time plays a critical role in how information is generated, retrieved, and interpreted. In this survey, we provide a comprehensive overview of Temporal Question Answering (TQA), a…

cs.CL2025

A Study into Investigating Temporal Robustness of LLMs

Jonas Wallat, Abdelrahman Abdallah, Adam Jatowt +1

Large Language Models (LLMs) encapsulate a surprising amount of factual world knowledge. However, their performance on temporal questions and historical knowledge is limited becaus…

cs.CL2025

From Retrieval to Generation: Comparing Different Approaches

Abdelrahman Abdallah, Jamshid Mozafari, Bhawna Piryani +2

Knowledge-intensive tasks, particularly open-domain question answering (ODQA), document reranking, and retrieval-augmented language modeling, require a balance between retrieval ac…