4 papers
Analyzing LLM Instruction Optimization for Tabular Fact Verification
Xiaotang Du, Giwon Hong, Wai-Chung Kwan +4
Instruction optimization provides a lightweight, model-agnostic approach to enhancing the reasoning performance of large language models (LLMs). This paper presents the first syste…
MGen: Millions of Naturally Occurring Generics in Context
Gustavo Cilleruelo, Emily Allaway, Barry Haddow +1
MGen is a dataset of over 4 million naturally occurring generic and quantified sentences extracted from diverse textual sources. Sentences in the dataset have long context document…
MatheMagic: Generating Dynamic Mathematics Benchmarks Robust to Memorization
Dayyán O'Brien, Barry Haddow, Emily Allaway +1
Conducting contamination-free evaluation of mathematical capabilities can be difficult for two reasons: models may memorize a test set once it is made public, and current mathemati…
Generics are puzzling. Can language models find the missing piece?
Gustavo Cilleruelo Calderón, Emily Allaway, Barry Haddow +1
Generic sentences express generalisations about the world without explicit quantification. Although generics are central to everyday communication, building a precise semantic fram…