Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Pitfalls of Evaluating Language Models with Open Benchmarks
Md. Najib Hasan, Md Mahadi Hassan Sibat, Mohammad Fakhruddin Babar +3
Open Large Language Model (LLM) benchmarks, such as HELM and BIG-Bench, provide standardized and transparent evaluation protocols that support comparative analysis, reproducibility…
cs.CL2024
LLMs as On-demand Customizable Service
Souvika Sarkar, Mohammad Fakhruddin Babar, Monowar Hasan +1
Large Language Models (LLMs) have demonstrated remarkable language understanding and generation capabilities. However, training, deploying, and accessing these models pose notable…