activity
20242026
collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Massively Multilingual Joint Segmentation and Glossing

Michael Ginn, Lindia Tjuatja, Enora Rice +5

Automated interlinear gloss prediction with neural networks is a promising approach to accelerate language documentation efforts. However, while state-of-the-art models like GlossL…

cs.CL2026

Neural Induction of Finite-State Transducers

Michael Ginn, Alexis Palmer, Mans Hulden

Finite-State Transducers (FSTs) are effective models for string-to-string rewriting tasks, often providing the efficiency necessary for high-performance applications, but construct…

cs.CL2026

Speculative Decoding and the Curse of Multilinguality

Nirajan Paudel, Michael Ginn, Luc De Nardi +1

Speculative decoding is a popular technique for large language model (LLM) inference, enabling faster generation by drafting multiple tokens with a smaller draft model. However, th…

cs.CL2025

Is linguistically-motivated data augmentation worth it?

Ray Groshan, Michael Ginn, Alexis Palmer

Data augmentation, a widely-employed technique for addressing data scarcity, involves generating synthetic data examples which are then used to augment available training data. Res…

cs.CL2024

Tree Transformers are an Ineffective Model of Syntactic Constituency

Michael Ginn

Linguists have long held that a key aspect of natural language syntax is the recursive organization of language units into constituent structures, and research has suggested that c…

cs.CL2024

GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text

Michael Ginn, Lindia Tjuatja, Taiqi He +4

Language documentation projects often involve the creation of annotated text in a format such as interlinear glossed text (IGT), which captures fine-grained morphosyntactic analyse…