10 papers
Tokenization with Factorized Subword Encoding
David Samuel, Lilja Øvrelid
In recent years, language models have become increasingly larger and more complex. However, the input representations for these models continue to rely on simple and greedy subword…
Building a Norwegian Lexical Resource for Medical Entity Recognition
Ildikó Pilán, Pål H. Brekke, Lilja Øvrelid
We present a large Norwegian lexical resource of categorized medical terms. The resource merges information from large medical databases, and contains over 77,000 unique entries, i…
One-to-X analogical reasoning on word embeddings: a case for diachronic armed conflict prediction from news texts
Andrey Kutuzov, Erik Velldal, Lilja Øvrelid
We extend the well-known word analogy task to a one-to-X formulation, including one-to-none cases, when no correct answer exists. The task is cast as a relation discovery problem a…
Sentiment analysis is not solved! Assessing and probing sentiment classification
Jeremy Barnes, Lilja Øvrelid, Erik Velldal
Neural methods for SA have led to quantitative improvements over previous approaches, but these advances are not always accompanied with a thorough analysis of the qualitative diff…
Probing Multilingual Sentence Representations With X-Probe
Vinit Ravishankar, Lilja Øvrelid, Erik Velldal
This paper extends the task of probing sentence representations for linguistic insight in a multilingual domain. In doing so, we make two contributions: first, we provide datasets…
Diachronic word embeddings and semantic shifts: a survey
Andrey Kutuzov, Lilja Øvrelid, Terrence Szymanski +1
Recent years have witnessed a surge of publications aimed at tracing temporal changes in lexical semantics using distributional methods, particularly prediction-based word embeddin…