18 papers
Question Difficulty Estimation for Large Language Models via Answer Plausibility Scoring
Jamshid Mozafari, Bhawna Piryani, Adam Jatowt
Estimating question difficulty is a critical component in evaluating and improving large language models (LLMs) for question answering (QA). Existing approaches often rely on reada…
Pretraining Exposure Explains Popularity Judgments in Large Language Models
Jamshid Mozafari, Bhawna Piryani, Adam Jatowt
Large language models (LLMs) exhibit systematic preferences for well-known entities, a phenomenon often attributed to popularity bias. However, the extent to which these preference…
Context Convergence Improves Answering Inferential Questions
Jamshid Mozafari, Bhawna Piryani, Adam Jatowt
While Large Language Models (LLMs) are widely used in open-domain Question Answering (QA), their ability to handle inferential questions-where answers must be derived rather than d…
It's High Time: A Survey of Temporal Question Answering
Bhawna Piryani, Abdelrahman Abdallah, Jamshid Mozafari +2
Time plays a critical role in how information is generated, retrieved, and interpreted. In this survey, we provide a comprehensive overview of Temporal Question Answering (TQA), a…
BracketRank: Large Language Model Document Ranking via Reasoning-based Competitive Elimination
Abdelrahman Abdallah, Mohammed Ali, Bhawna Piryani +1
Reasoning-intensive retrieval requires deep semantic inference beyond surface-level keyword matching, posing a challenge for current LLM-based rerankers limited by context constrai…
How often do Answers Change? Estimating Recency Requirements in Question Answering
Bhawna Piryani, Zehra Mert, Adam Jatowt
Large language models (LLMs) often rely on outdated knowledge when answering time-sensitive questions, leading to confident yet incorrect responses. Without explicit signals indica…