1 paper
Do Xuan Long, Hai Nguyen Ngoc, Tiviatis Sim +5
We present the first systematic evaluation examining format bias in performance of large language models (LLMs). Our approach distinguishes between two categories of an evaluation…