activity
20202026
most citedRobust Text Line Detection in Historical Documents: Learning and Evaluation Methods

38 citations · 83 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition

Mélodie Boillet, Solène Tarride, Christopher Kermorvant

Benchmarks that reflect the diversity and complexity of real-world documents are essential for accurately evaluating Automatic Text Recognition (ATR) systems, especially Vision-Lar…

cs.CV2024★ 1 cited

Improving Automatic Text Recognition with Language Models in the PyLaia Open-Source Library

Solène Tarride, Yoann Schneider, Marie Generali-Lince +3

PyLaia is one of the most popular open-source software for Automatic Text Recognition (ATR), delivering strong performance in terms of speed and accuracy. In this paper, we outline…

cs.CV2024

The Socface Project: Large-Scale Collection, Processing, and Analysis of a Century of French Censuses

Mélodie Boillet, Solène Tarride, Manon Blanco +5

This paper presents a complete processing workflow for extracting information from French census lists from 1836 to 1936. These lists contain information about individuals living i…

cs.CV2023

Handwritten Text Recognition from Crowdsourced Annotations

Solène Tarride, Tristan Faine, Mélodie Boillet +2

In this paper, we explore different ways of training a model for handwritten text recognition when multiple imperfect or noisy transcriptions are available. We consider various tra…

cs.CV2023★ 14 cited

Large Scale Genealogical Information Extraction From Handwritten Quebec Parish Records

Solène Tarride, Martin Maarand, Mélodie Boillet +4

This paper presents a complete workflow designed for extracting information from Quebec handwritten parish registers. The acts in these documents contain individual and family info…

cs.CV2023

SIMARA: a database for key-value information extraction from full pages

Solène Tarride, Mélodie Boillet, Jean-François Moufflet +1

We propose a new database for information extraction from historical handwritten documents. The corpus includes 5,393 finding aids from six different series, dating from the 18th-2…