2 papers
cs.CL2026
Pitfalls of Evaluating Language Models with Open Benchmarks
Md. Najib Hasan, Md Mahadi Hassan Sibat, Mohammad Fakhruddin Babar +3
Open Large Language Model (LLM) benchmarks, such as HELM and BIG-Bench, provide standardized and transparent evaluation protocols that support comparative analysis, reproducibility…
cs.CL2025
Large Language Models for IT Automation Tasks: Are We There Yet?
Md Mahadi Hassan, John Salvador, Akond Rahman +1
LLMs show promise in code generation, yet their effectiveness for IT automation tasks, particularly for tools like Ansible, remains understudied. Existing benchmarks rely primarily…