563 citations
- Sorbonne UniversitéFR2 papers
- Asus (Taiwan)TW1 paper
- Birla Institute of Technology and Science, PilaniIN1 paper
- Booz Allen Hamilton (United States)US1 paper
- Brown UniversityUS1 paper
- Center PointUS1 paper
- Charles River Analytics (United States)US1 paper
- CyberOptics (United States)US1 paper
- École des hautes études en sciences socialesFR1 paper
- École Nationale des ChartesFR1 paper
- École Normale Supérieure - PSLFR1 paper
- Histoire et Sources des Mondes AntiquesFR1 paper
8 papers
MANTa: Efficient Gradient-Based Tokenization for Robust End-to-End Language Modeling
Nathan Godey, Roman Castagné, Éric de la Clergerie +1
Static subword tokenization algorithms have been an essential component of recent works on language modeling. However, their static nature results in important flaws that degrade t…
Technological taxonomies for hypernym and hyponym retrieval in patent texts
You Zuo, Yixuan Li, Alma Parias García +1
This paper presents an automatic approach to creating taxonomies of technical terms based on the Cooperative Patent Classification (CPC). The resulting taxonomy contains about 170k…
You Actually Look Twice At it (YALTAi): using an object detection approach instead of region segmentation within the Kraken engine
Thibault Clérice
Layout Analysis (the identification of zones and their classification) is the first step along line segmentation in Optical Character Recognition and similar tasks. The ability of…
DP-Parse: Finding Word Boundaries from Raw Speech with an Instance Lexicon
Robin Algayres, Tristan Ricoul, Julien Karadayi +5
Finding word boundaries in continuous speech is challenging as there is little or no equivalent of a 'space' delimiter between words. Popular Bayesian non-parametric models for tex…
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
Arij Riabi, Benoît Sagot, Djamé Seddah
Recent impressive improvements in NLP, largely based on the success of contextual neural language models, have been mostly demonstrated on at most a couple dozen high-resource lang…
Multitask Prompted Training Enables Zero-Shot Task Generalization
Victor Sanh, Albert Webson, Colin Raffel +38
Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a…