3 papers
cs.IR2026
Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark
Nishan Pantha, Pranath Reddy Kumbam, Sajil Awale +8
Scientists increasingly rely on open-source tools to support their research workflows, yet discovering relevant software among over 600 million GitHub repositories remains challeng…
cs.CL2026
RECOM: A Validity Discrimination Tradeoff in Automatic Metrics for Open Ended Reddit Question Answering
Pushwitha Krishnappa, Amit Das, Vinija Jain +2
Automatic metrics are the default for evaluating LLM-generated text, yet a metric is quietly asked to do two jobs: tell genuine content alignment from surface coincidence (validity…
cs.CL2026
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
Pushwitha Krishnappa, Amit Das, Vinija Jain +2
Large Language Models (LLMs) are increasingly deployed for open-domain question answering, yet their alignment with human perspectives on temporally recent information remains unde…