Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024
Large-scale cloze evaluation reveals that token prediction tasks are neither lexically nor semantically aligned
Cassandra L. Jacobs, Loïc Grobol, Alvin Tsang
In this work we compare the generative behavior at the next token prediction level in several language models by comparing them to human productions in the cloze task. We find that…
cs.CL2023
Incorporating Annotator Uncertainty into Representations of Discourse Relations
S. Magalí López Cortez, Cassandra L. Jacobs
Annotation of discourse relations is a known difficult task, especially for non-expert annotators. In this paper, we investigate novice annotators' uncertainty on the annotation of…
cs.CL2022
Lost in Space Marking
Cassandra L. Jacobs, Yuval Pinter
We look at a decision taken early in training a subword tokenizer, namely whether it should be the word-initial token that carries a special mark, or the word-final one. Based on s…