4 papers
RespondeoQA: a Benchmark for Bilingual Latin-English Question Answering
Marisa Hudspeth, Patrick J. Burns, Brendan O'Connor
We introduce a benchmark dataset for question answering and translation in bilingual Latin and English settings, containing about 7,800 question-answer pairs. The questions are dra…
Contextual morphologically-guided tokenization for Latin encoder models
Marisa Hudspeth, Patrick J. Burns, Brendan O'Connor
Tokenization is a critical component of language model pretraining, yet standard tokenization methods often prioritize information-theoretical goals like high compression and low f…
BAP v2: An Enhanced Task Framework for Instruction Following in Minecraft Dialogues
Prashant Jayannavar, Liliang Ren, Marisa Hudspeth +6
Developing interactive agents that can understand language, perceive their surroundings, and act within the physical world is a long-standing goal of AI research. The Minecraft Col…
Evaluating Morphological Alignment of Tokenizers in 70 Languages
Catherine Arnett, Marisa Hudspeth, Brendan O'Connor
While tokenization is a key step in language modeling, with effects on model training and performance, it remains unclear how to effectively evaluate tokenizer quality. One propose…