1 paper
Alhasan Mahmood, Samir Abdaljalil, Hasan Kurban
Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's language can invert backbone rankings. We loca…