19 citations · 29 across the 8 of their papers we have counts for
12 papers · 1 filter
Do Multilingual Language Models Think Better in English?
Julen Etxaniz, Gorka Azkune, Aitor Soroa +2
Translate-test is a popular technique to improve the performance of multilingual language models. This approach works by translating the input into English using an external machin…
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang +51
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…
Noisy Channel for Automatic Text Simplification
Oscar M Cumbicus-Pineda, Iker Gutiérrez-Fandiño, Itziar Gonzalez-Dios +1
In this paper we present a simple re-ranking method for Automatic Sentence Simplification based on the noisy channel scheme. Instead of directly computing the best simplification g…
Documenting Geographically and Contextually Diverse Data Sources: The BigScience Catalogue of Language Data and Resources
Angelina McMillan-Major, Zaid Alyafeai, Stella Biderman +15
In recent years, large-scale data collection efforts have prioritized the amount of data collected in order to improve the modeling capabilities of large language models. This prio…
Improving Conversational Question Answering Systems after Deployment using Feedback-Weighted Learning
Jon Ander Campos, Kyunghyun Cho, Arantxa Otegi +3
The interaction of conversational systems with users poses an exciting opportunity for improving them after deployment, but little evidence has been provided of its feasibility. In…
DoQA -- Accessing Domain-Specific FAQs via Conversational QA
Jon Ander Campos, Arantxa Otegi, Aitor Soroa +3
The goal of this work is to build conversational Question Answering (QA) interfaces for the large body of domain-specific information available in FAQ sites. We present DoQA, a dat…