7 papers
Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing
Arthur Amalvy, Vincent Labatut, Xavier Bost +1
While annotated corpora are crucial in the field of natural language processing (NLP), those containing copyrighted material are difficult to exchange among researchers. Yet, such…
Beyond Known Facts: Generating Unseen Temporal Knowledge to Address Data Contamination in LLM Evaluation
Arthur Amalvy, Hen-Hsen Huang
The automatic extraction of information is important for populating large web knowledge bases such as Wikidata. The temporal version of that task, temporal knowledge graph extracti…
Persistent Homology of Topic Networks for the Prediction of Reader Curiosity
Manuel D. S. Hopp, Vincent Labatut, Arthur Amalvy +4
Reader curiosity, the drive to seek information, is crucial for textual engagement, yet remains relatively underexplored in NLP. Building on Loewenstein's Information Gap Theory, w…
The Role of Natural Language Processing Tasks in Automatic Literary Character Network Construction
Arthur Amalvy, Vincent Labatut, Richard Dufour
The automatic extraction of character networks from literary texts is generally carried out using natural language processing (NLP) cascading pipelines. While this approach is wide…
Interconnected Kingdoms: Comparing 'A Song of Ice and Fire' Adaptations Across Media Using Complex Networks
Arthur Amalvy, Madeleine Janickyj, Shane Mannion +2
In this article, we propose and apply a method to compare adaptations of the same story across different media. We tackle this task by modelling such adaptations through character…
Annotation Guidelines for Corpus Novelties: Part 1 -- Named Entity Recognition
Arthur Amalvy, Vincent Labatut
The Novelties corpus is a collection of novels (and parts of novels) annotated for Named Entity Recognition (NER) among other tasks. This document describes the guidelines applied…