activity
20212023
most citedThe Web Is Your Oyster - Knowledge-Intensive NLP against a Very Large Web Corpus

24 citations · 25 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL20231 cited

FinGPT: Large Generative Models for a Small Language

Risto Luukkonen, Ville Komulainen, Jouni Luoma +18

Large language models (LLMs) excel in many tasks in NLP and beyond, but most open models have very limited coverage of smaller languages and LLM work tends to focus on languages wh…

cs.CL2023

GAIA Search: Hugging Face and Pyserini Interoperability for NLP Training Data Exploration

Aleksandra Piktus, Odunayo Ogundepo, Christopher Akiki +6

Noticing the urgent need to provide tools for fast and user-friendly qualitative analysis of large-scale textual corpora of the modern NLP, we propose to turn to the mature and wel…

cs.CL2023

The ROOTS Search Tool: Data Transparency for LLMs

Aleksandra Piktus, Christopher Akiki, Paulo Villegas +5

ROOTS is a 1.6TB multilingual text corpus developed for the training of BLOOM, currently the largest language model explicitly accompanied by commensurate data governance efforts.…

cs.IR2022

Improving Wikipedia Verifiability with AI

Fabio Petroni, Samuel Broscheit, Aleksandra Piktus +10

Verifiability is a core content policy of Wikipedia: claims that are likely to be challenged need to be backed by citations. There are millions of articles available online and tho…

cs.CL202124 cited

The Web Is Your Oyster - Knowledge-Intensive NLP against a Very Large Web Corpus

Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin +8

In order to address increasing demands of real-world applications, the research for knowledge-intensive NLP (KI-NLP) should advance by capturing the challenges of a truly open-doma…