activity
20202023
most citedNoCoLA: The Norwegian Corpus of Linguistic Acceptability

4 citations · 7 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL20234 cited

NoCoLA: The Norwegian Corpus of Linguistic Acceptability

Matias Jentoft, David Samuel

While there has been a surge of large language models for Norwegian in recent years, we lack any tool to evaluate their understanding of grammaticality. We present two new Norwegia…

cs.CL2023

Tokenization with Factorized Subword Encoding

David Samuel, Lilja Øvrelid

In recent years, language models have become increasingly larger and more complex. However, the input representations for these models continue to rely on simple and greedy subword…

cs.CL2022

EventGraph at CASE 2021 Task 1: A General Graph-based Approach to Protest Event Extraction

Huiling You, David Samuel, Samia Touileb +1

This paper presents our submission to the 2022 edition of the CASE 2021 shared task 1, subtask 4. The EventGraph system adapts an end-to-end, graph-based semantic parser to the tas…

cs.CL20221 cited

EventGraph: Event Extraction as Semantic Graph Parsing

Huiling You, David Samuel, Samia Touileb +1

Event extraction involves the detection and extraction of both the event triggers and corresponding event arguments. Existing systems often decompose event extraction into multiple…

cs.CL2022

Direct parsing to sentiment graphs

David Samuel, Jeremy Barnes, Robin Kurtz +3

This paper demonstrates how a graph-based semantic parser can be applied to the task of structured sentiment analysis, directly predicting sentiment graphs from text. We advance th…

cs.CL2021

RobeCzech: Czech RoBERTa, a monolingual contextualized language representation model

Milan Straka, Jakub Náplava, Jana Straková +1

We present RobeCzech, a monolingual RoBERTa language representation model trained on Czech data. RoBERTa is a robustly optimized Transformer-based pretraining approach. We show tha…