1 citations · 1 across the 2 of their papers we have counts for
3 papers
MatheMagic: Generating Dynamic Mathematics Benchmarks Robust to Memorization
Dayyán O'Brien, Barry Haddow, Emily Allaway +1
Conducting contamination-free evaluation of mathematical capabilities can be difficult for two reasons: models may memorize a test set once it is made public, and current mathemati…
MGen: Millions of Naturally Occurring Generics in Context
Gustavo Cilleruelo, Emily Allaway, Barry Haddow +1
MGen is a dataset of over 4 million naturally occurring generic and quantified sentences extracted from diverse textual sources. Sentences in the dataset have long context document…
Generics are puzzling. Can language models find the missing piece?
Gustavo Cilleruelo Calderón, Emily Allaway, Barry Haddow +1
Generic sentences express generalisations about the world without explicit quantification. Although generics are central to everyday communication, building a precise semantic fram…