1 paper
Nathan A. Chi, Teodor Malchev, Riley Kong +5
We introduce modeLing, a novel benchmark of Linguistics Olympiad-style puzzles which tests few-shot reasoning in AI systems. Solving these puzzles necessitates inferring aspects of…