1 paper
Sofie Helene Bruun, Dan Saattrup Smart
We create high-quality datasets for LLM evaluation of logical reasoning skills across nine different languages, which have been manually checked by fluent speakers. The datasets co…