collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

RespondeoQA: a Benchmark for Bilingual Latin-English Question Answering

Marisa Hudspeth, Patrick J. Burns, Brendan O'Connor

We introduce a benchmark dataset for question answering and translation in bilingual Latin and English settings, containing about 7,800 question-answer pairs. The questions are dra…

cs.CL2025

Contextual morphologically-guided tokenization for Latin encoder models

Marisa Hudspeth, Patrick J. Burns, Brendan O'Connor

Tokenization is a critical component of language model pretraining, yet standard tokenization methods often prioritize information-theoretical goals like high compression and low f…

cs.CL2025

BAP v2: An Enhanced Task Framework for Instruction Following in Minecraft Dialogues

Prashant Jayannavar, Liliang Ren, Marisa Hudspeth +6

Developing interactive agents that can understand language, perceive their surroundings, and act within the physical world is a long-standing goal of AI research. The Minecraft Col…

cs.CL2025

Evaluating Morphological Alignment of Tokenizers in 70 Languages

Catherine Arnett, Marisa Hudspeth, Brendan O'Connor

While tokenization is a key step in language modeling, with effects on model training and performance, it remains unclear how to effectively evaluate tokenizer quality. One propose…

cs.CL2024

Latin Treebanks in Review: An Evaluation of Morphological Tagging Across Time

Marisa Hudspeth, Brendan O'Connor, Laure Thompson

Existing Latin treebanks draw from Latin's long written tradition, spanning 17 centuries and a variety of cultures. Recent efforts have begun to harmonize these treebanks' annotati…