2 citations · 3 across the 5 of their papers we have counts for
6 papers · 1 filter
Revising RVL-CDIP: Quantifying Errors and Test-Train Overlap
Stefan Larson, Attila Nagy, Sam Desai +8
RVL-CDIP is a popular dataset for benchmarking document classifiers. However, the dataset contains ample amounts of label errors as well as non-trivial amounts of test-train overla…
From News to Summaries: Building a Hungarian Corpus for Extractive and Abstractive Summarization
Botond Barta, Dorina Lakatos, Attila Nagy +2
Training summarization models requires substantial amounts of training data. However for less resourceful languages like Hungarian, openly available models and datasets are notably…
TreeSwap: Data Augmentation for Machine Translation via Dependency Subtree Swapping
Attila Nagy, Dorina Lakatos, Botond Barta +1
Data augmentation methods for neural machine translation are particularly useful when limited amount of training data is available, which is often the case when dealing with low-re…
Data Augmentation for Machine Translation via Dependency Subtree Swapping
Attila Nagy, Dorina Petra Lakatos, Botond Barta +2
We present a generic framework for data augmentation via dependency subtree swapping that is applicable to machine translation. We extract corresponding subtrees from the dependenc…
HunSum-1: an Abstractive Summarization Dataset for Hungarian
Botond Barta, Dorina Lakatos, Attila Nagy +2
We introduce HunSum-1: a dataset for Hungarian abstractive summarization, consisting of 1.14M news articles. The dataset is built by collecting, cleaning and deduplicating data fro…
Syntax-based data augmentation for Hungarian-English machine translation
Attila Nagy, Patrick Nanys, Balázs Frey Konrád +2
We train Transformer-based neural machine translation models for Hungarian-English and English-Hungarian using the Hunglish2 corpus. Our best models achieve a BLEU score of 40.0 on…