2 papers
cs.CL2026
Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark
Yu Wu, Ke Shu, Jonas Fischer +4
This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the p…
cs.LG2025
ARETE: an R package for Automated REtrieval from TExt with large language models
Vasco V. Branco, Jandó Benedek, Lidia Pivovarova +2
1. A hard stop for the implementation of rigorous conservation initiatives is our lack of key species data, especially occurrence data. Furthermore, researchers have to contend wit…