activity
20152023
most citedStudying the Wikipedia Hyperlink Graph for Relatedness and Disambiguation

19 citations · 29 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL20232 cited

Do Multilingual Language Models Think Better in English?

Julen Etxaniz, Gorka Azkune, Aitor Soroa +2

Translate-test is a popular technique to improve the performance of multilingual language models. This approach works by translating the input into English using an external machin…

cs.CL202365 cited

The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

Hugo Laurençon, Lucile Saulnier, Thomas Wang +51

As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop,…

cs.CL2022

Noisy Channel for Automatic Text Simplification

Oscar M Cumbicus-Pineda, Iker Gutiérrez-Fandiño, Itziar Gonzalez-Dios +1

In this paper we present a simple re-ranking method for Automatic Sentence Simplification based on the noisy channel scheme. Instead of directly computing the best simplification g…

cs.CL20222 cited

Documenting Geographically and Contextually Diverse Data Sources: The BigScience Catalogue of Language Data and Resources

Angelina McMillan-Major, Zaid Alyafeai, Stella Biderman +15

In recent years, large-scale data collection efforts have prioritized the amount of data collected in order to improve the modeling capabilities of large language models. This prio…

cs.CL2020

Improving Conversational Question Answering Systems after Deployment using Feedback-Weighted Learning

Jon Ander Campos, Kyunghyun Cho, Arantxa Otegi +3

The interaction of conversational systems with users poses an exciting opportunity for improving them after deployment, but little evidence has been provided of its feasibility. In…

cs.CL2020

DoQA -- Accessing Domain-Specific FAQs via Conversational QA

Jon Ander Campos, Arantxa Otegi, Aitor Soroa +3

The goal of this work is to build conversational Question Answering (QA) interfaces for the large body of domain-specific information available in FAQ sites. We present DoQA, a dat…