7 citations · 15 across the 4 of their papers we have counts for
6 papers
Dynaword: From One-shot to Continuously Developed Datasets
Kenneth Enevoldsen, Kristian Nørgaard Jensen, Jan Kostkan +14
Large-scale datasets are foundational for research and development in natural language processing. However, current approaches face three key challenges: (1) reliance on ambiguousl…
Skill Extraction from Job Postings using Weak Supervision
Mike Zhang, Kristian Nørgaard Jensen, Rob van der Goot +1
Aggregated data obtained from job postings provide powerful insights into labor market demands, and emerging skills, and aid job matching. However, most extraction approaches are s…
Kompetencer: Fine-grained Skill Classification in Danish Job Postings via Distant Supervision and Transfer Learning
Mike Zhang, Kristian Nørgaard Jensen, Barbara Plank
Skill Classification (SC) is the task of classifying job competences from job postings. This work is the first in SC applied to Danish job vacancy data. We release the first Danish…
SkillSpan: Hard and Soft Skill Extraction from English Job Postings
Mike Zhang, Kristian Nørgaard Jensen, Sif Dam Sonniks +1
Skill Extraction (SE) is an important and widely-studied task useful to gain insights into labor market dynamics. However, there is a lacuna of datasets and annotation guidelines;…
DaN+: Danish Nested Named Entities and Lexical Normalization
Barbara Plank, Kristian Nørgaard Jensen, Rob van der Goot
This paper introduces DaN+, a new multi-domain corpus and annotation guidelines for Danish nested named entities (NEs) and lexical normalization to support research on cross-lingua…
De-identification of Privacy-related Entities in Job Postings
Kristian Nørgaard Jensen, Mike Zhang, Barbara Plank
De-identification is the task of detecting privacy-related entities in text, such as person names, emails and contact data. It has been well-studied within the medical domain. The…