papers

Publications (13)

cs.CL2021

Supersense and Sensibility: Proxy Tasks for Semantic Annotation of Prepositions

Luke Gessler, Shira Wein, Nathan Schneider

Prepositional supersense annotation is time-consuming and requires expert training. Here, we present two sensible methods for obtaining prepositional supersense annotations by elic…

cs.CL2024

TAMS: Translation-Assisted Morphological Segmentation

Enora Rice, Ali Marashian, Luke Gessler +2

Canonical morphological segmentation is the process of analyzing words into the standard (aka underlying) forms of their constituent morphemes. This is a core task in language docu…

cs.CL2023

Syntactic Inductive Bias in Transformer Language Models: Especially Helpful for Low-Resource Languages?

Luke Gessler, Nathan Schneider

A line of work on Transformer-based language models such as BERT has attempted to use syntactic inductive bias to enhance the pretraining process, on the theory that building synta…

cs.CL2021

BERT Has Uncommon Sense: Similarity Ranking for Word Sense BERTology

Luke Gessler, Nathan Schneider

An important question concerning contextualized word embedding (CWE) models like BERT is how well they can represent different word senses, especially those in the long tail of unc…

cs.CL2020

A Summary of the First Workshop on Language Technology for Language Documentation and Revitalization

Graham Neubig, Shruti Rijhwani, Alexis Palmer +21

Despite recent advances in natural language processing and other language technology, the application of such technology to language documentation and conservation has been limited…

cs.CL2020

Supervised Grapheme-to-Phoneme Conversion of Orthographic Schwas in Hindi and Punjabi

Aryaman Arora, Luke Gessler, Nathan Schneider

Hindi grapheme-to-phoneme (G2P) conversion is mostly trivial, with one exception: whether a schwa represented in the orthography is pronounced or unpronounced (deleted). Previous w…

cs.CL2020

AMALGUM -- A Free, Balanced, Multilayer English Web Corpus

Luke Gessler, Siyao Peng, Yang Liu +3

We present a freely available, genre-balanced English web corpus totaling 4M tokens and featuring a large number of high-quality automatic annotation layers, including dependency t…

cs.CL2025

From Priest to Doctor: Domain Adaptation for Low-Resource Neural Machine Translation

Ali Marashian, Enora Rice, Luke Gessler +2

Many of the world's languages have insufficient data to train high-performing general neural machine translation (NMT) models, let alone domain-specific models, and often the only…

cs.CL2021

DisCoDisCo at the DISRPT2021 Shared Task: A System for Discourse Segmentation, Classification, and Connective Detection

Luke Gessler, Shabnam Behzad, Yang Janet Liu +3

This paper describes our submission to the DISRPT2021 Shared Task on Discourse Unit Segmentation, Connective Detection, and Relation Classification. Our system, called DisCoDisCo,…

cs.CL2024

PrOnto: Language Model Evaluations for 859 Languages

Luke Gessler

Evaluation datasets are critical resources for measuring the quality of pretrained language models. However, due to the high cost of dataset annotation, these resources are scarce…

cs.CL2023

MicroBERT: Effective Training of Low-resource Monolingual BERTs through Parameter Reduction and Multitask Learning

Luke Gessler, Amir Zeldes

Transformer language models (TLMs) are critical for most NLP tasks, but they are difficult to create for low-resource languages because of how much pretraining data they require. I…

cs.CL2024

eRST: A Signaled Graph Theory of Discourse Relations and Organization

Amir Zeldes, Tatsuya Aoyama, Yang Janet Liu +3

In this article we present Enhanced Rhetorical Structure Theory (eRST), a new theoretical framework for computational discourse analysis, based on an expansion of Rhetorical Struct…

cs.CL2023

GENTLE: A Genre-Diverse Multilayer Challenge Set for English NLP and Linguistic Evaluation

Tatsuya Aoyama, Shabnam Behzad, Luke Gessler +6

We present GENTLE, a new mixed-genre English challenge corpus totaling 17K tokens and consisting of 8 unusual text types for out-of domain evaluation: dictionary entries, esports c…