4 papers
The Validity of Coreference-based Evaluations of Natural Language Understanding
Ian Porada
In this thesis, I refine our understanding as to what conclusions we can reach from coreference-based evaluations by expanding existing evaluation practices and considering the ext…
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
Ian Porada, Jackie Chi Kit Cheung
Challenge sets such as the Winograd Schema Challenge (WSC) are used to benchmark systems' ability to resolve ambiguities in natural language. If one assumes as in existing work tha…
Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective
Ian Porada, Alexandra Olteanu, Kaheer Suleman +2
It is increasingly common to evaluate the same coreference resolution (CR) model on multiple datasets. Do these multi-dataset evaluations allow us to draw meaningful conclusions ab…
A Controlled Reevaluation of Coreference Resolution Models
Ian Porada, Xiyuan Zou, Jackie Chi Kit Cheung
All state-of-the-art coreference resolution (CR) models involve finetuning a pretrained language model. Whether the superior performance of one CR model over another is due to the…