1 citations · 2 across the 4 of their papers we have counts for
4 papers
NLPre: a revised approach towards language-centric benchmarking of Natural Language Preprocessing systems
Martyna Wiącek, Piotr Rybak, Łukasz Pszenny +1
With the advancements of transformer-based architectures, we observe the rise of natural language preprocessing (NLPre) tools capable of solving preliminary NLP tasks (e.g. tokenis…
Transferring BERT Capabilities from High-Resource to Low-Resource Languages Using Vocabulary Matching
Piotr Rybak
Pre-trained language models have revolutionized the natural language understanding landscape, most notably BERT (Bidirectional Encoder Representations from Transformers). However,…
MAUPQA: Massive Automatically-created Polish Question Answering Dataset
Piotr Rybak
Recently, open-domain question answering systems have begun to rely heavily on annotated datasets to train neural passage retrievers. However, manually annotating such datasets is…
Going beyond research datasets: Novel intent discovery in the industry setting
Aleksandra Chrabrowa, Tsimur Hadeliya, Dariusz Kajtoch +2
Novel intent discovery automates the process of grouping similar messages (questions) to identify previously unknown intents. However, current research focuses on publicly availabl…