Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Pitfalls of Evaluating Language Models with Open Benchmarks
Md. Najib Hasan, Md Mahadi Hassan Sibat, Mohammad Fakhruddin Babar +3
Open Large Language Model (LLM) benchmarks, such as HELM and BIG-Bench, provide standardized and transparent evaluation protocols that support comparative analysis, reproducibility…
cs.CL2025
Benchmarking LLMs on the Semantic Overlap Summarization Task
John Salvador, Naman Bansal, Mousumi Akter +3
Semantic Overlap Summarization (SOS) is a constrained multi-document summarization task, where the constraint is to capture the common/overlapping information between two alternati…