2 papers
cs.CL2026
Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks
Fahmid Shahriar Iqbal, Ritam Dutt, Soumitra Das +2
Time and event expression extraction are fundamental temporal reasoning tasks, but the problem remains difficult due to annotation ambiguity, domain sensitivity, and unstable model…
cs.CL2026
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
Rafid Ishrak Jahan, Fahmid Shahriar Iqbal, Sagnik Ray Choudhury
Long-form question answering (LFQA) demands nuanced evaluation of multi-sentence explanatory responses, yet existing metrics often fail to reflect human judgment. We present LFQA-H…