9 citations · 40 across the 19 of their papers we have counts for
29 papers · 1 filter
Stop Measuring Calibration When Humans Disagree
Joris Baan, Wilker Aziz, Barbara Plank +1
Calibration is a popular framework to evaluate whether a classifier knows when it does not know - i.e., its predictive probabilities are a good indication of how likely a predictio…
The 'Problem' of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation
Barbara Plank
Human variation in labeling is often considered noise. Annotation projects for machine learning (ML) aim at minimizing human label variation, with the assumption to maximize data q…
Spectral Probing
Max Müller-Eberstein, Rob van der Goot, Barbara Plank
Linguistic information is encoded at varying timescales (subwords, phrases, etc.) and communicative levels, such as syntax and semantics. Contextualized embeddings have analogously…
Evidence > Intuition: Transferability Estimation for Encoder Selection
Elisa Bassignana, Max Müller-Eberstein, Mike Zhang +1
With the increase in availability of large pre-trained language models (LMs) in Natural Language Processing (NLP), it becomes critical to assess their fit for a specific target tas…
CrossRE: A Cross-Domain Dataset for Relation Extraction
Elisa Bassignana, Barbara Plank
Relation Extraction (RE) has attracted increasing attention, but current RE evaluation is limited to in-domain evaluation setups. Little is known on how well a RE system fares in c…
Skill Extraction from Job Postings using Weak Supervision
Mike Zhang, Kristian Nørgaard Jensen, Rob van der Goot +1
Aggregated data obtained from job postings provide powerful insights into labor market demands, and emerging skills, and aid job matching. However, most extraction approaches are s…