1 paper
Aymene Berriche, Cathrine Shalby, Mohannad Alhanahnah +1
Large language model (LLM) benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cyberse…