activity
20222024
most citedLarge Scale Genealogical Information Extraction From Handwritten Quebec Parish Records

14 citations · 15 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV20241 cited

Improving Automatic Text Recognition with Language Models in the PyLaia Open-Source Library

Solène Tarride, Yoann Schneider, Marie Generali-Lince +3

PyLaia is one of the most popular open-source software for Automatic Text Recognition (ATR), delivering strong performance in terms of speed and accuracy. In this paper, we outline…

cs.CV202314 cited

Large Scale Genealogical Information Extraction From Handwritten Quebec Parish Records

Solène Tarride, Martin Maarand, Mélodie Boillet +4

This paper presents a complete workflow designed for extracting information from Quebec handwritten parish registers. The acts in these documents contain individual and family info…

cs.CV2023

SIMARA: a database for key-value information extraction from full pages

Solène Tarride, Mélodie Boillet, Jean-François Moufflet +1

We propose a new database for information extraction from historical handwritten documents. The corpus includes 5,393 finding aids from six different series, dating from the 18th-2…

cs.CV2023

Key-value information extraction from full handwritten pages

Solène Tarride, Mélodie Boillet, Christopher Kermorvant

We propose a Transformer-based approach for information extraction from digitized handwritten documents. Our approach combines, in a single model, the different steps that were so…

cs.CV2023

Détection d'Objets dans les documents numérisés par réseaux de neurones profonds

Mélodie Boillet

In this thesis, we study multiple tasks related to document layout analysis such as the detection of text lines, the splitting into acts or the detection of the writing support. Th…

cs.CV2022

Confidence Estimation for Object Detection in Document Images

Mélodie Boillet, Christopher Kermorvant, Thierry Paquet

Deep neural networks are becoming increasingly powerful and large and always require more labelled data to be trained. However, since annotating data is time-consuming, it is now n…