output
20192022
most citedMultitask Prompted Training Enables Zero-Shot Task Generalization

563 citations

8 papers

cs.CL2022

MANTa: Efficient Gradient-Based Tokenization for Robust End-to-End Language Modeling

Nathan Godey, Roman Castagné, Éric de la Clergerie +1

Static subword tokenization algorithms have been an essential component of recent works on language modeling. However, their static nature results in important flaws that degrade t…

cs.CL2022

Technological taxonomies for hypernym and hyponym retrieval in patent texts

You Zuo, Yixuan Li, Alma Parias García +1

This paper presents an automatic approach to creating taxonomies of technical terms based on the Cooperative Patent Classification (CPC). The resulting taxonomy contains about 170k…

cs.CV2022★ 2 cited

You Actually Look Twice At it (YALTAi): using an object detection approach instead of region segmentation within the Kraken engine

Thibault Clérice

Layout Analysis (the identification of zones and their classification) is the first step along line segmentation in Optical Character Recognition and similar tasks. The ability of…

cs.CL2022★ 3 cited

DP-Parse: Finding Word Boundaries from Raw Speech with an Instance Lexicon

Robin Algayres, Tristan Ricoul, Julien Karadayi +5

Finding word boundaries in continuous speech is challenging as there is little or no equivalent of a 'space' delimiter between words. Popular Bayesian non-parametric models for tex…

cs.CL2021★ 1 cited

Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?

Arij Riabi, Benoît Sagot, Djamé Seddah

Recent impressive improvements in NLP, largely based on the success of contextual neural language models, have been mostly demonstrated on at most a couple dozen high-resource lang…

cs.LG2021★ 563 cited

Multitask Prompted Training Enables Zero-Shot Task Generalization

Victor Sanh, Albert Webson, Colin Raffel +38

Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a…