Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
RECOM: A Validity Discrimination Tradeoff in Automatic Metrics for Open Ended Reddit Question Answering
Pushwitha Krishnappa, Amit Das, Vinija Jain +2
Automatic metrics are the default for evaluating LLM-generated text, yet a metric is quietly asked to do two jobs: tell genuine content alignment from surface coincidence (validity…
cs.CL2026
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
Pushwitha Krishnappa, Amit Das, Vinija Jain +2
Large Language Models (LLMs) are increasingly deployed for open-domain question answering, yet their alignment with human perspectives on temporally recent information remains unde…