2 papers
cs.CL2024
Remember This Event That Year? Assessing Temporal Information and Reasoning in Large Language Models
Himanshu Beniwal, Dishant Patel, Kowsik Nandagopan D +3
Large Language Models (LLMs) are increasingly ubiquitous, yet their ability to retain and reason about temporal information remains limited, hindering their application in real-wor…
cs.CL2024
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
Ankit Yadav, Himanshu Beniwal, Mayank Singh
Driven by the surge in code generation using large language models (LLMs), numerous benchmarks have emerged to evaluate these LLMs capabilities. We conducted a large-scale human ev…