activity
20222026
most citedHunSum-1: an Abstractive Summarization Dataset for Hungarian

2 citations · 3 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

Revising RVL-CDIP: Quantifying Errors and Test-Train Overlap

Stefan Larson, Attila Nagy, Sam Desai +8

RVL-CDIP is a popular dataset for benchmarking document classifiers. However, the dataset contains ample amounts of label errors as well as non-trivial amounts of test-train overla…

cs.CL2024

From News to Summaries: Building a Hungarian Corpus for Extractive and Abstractive Summarization

Botond Barta, Dorina Lakatos, Attila Nagy +2

Training summarization models requires substantial amounts of training data. However for less resourceful languages like Hungarian, openly available models and datasets are notably…

cs.CL2023

TreeSwap: Data Augmentation for Machine Translation via Dependency Subtree Swapping

Attila Nagy, Dorina Lakatos, Botond Barta +1

Data augmentation methods for neural machine translation are particularly useful when limited amount of training data is available, which is often the case when dealing with low-re…

cs.CL2023★ 1 cited

Data Augmentation for Machine Translation via Dependency Subtree Swapping

Attila Nagy, Dorina Petra Lakatos, Botond Barta +2

We present a generic framework for data augmentation via dependency subtree swapping that is applicable to machine translation. We extract corresponding subtrees from the dependenc…

cs.CL2023★ 2 cited

HunSum-1: an Abstractive Summarization Dataset for Hungarian

Botond Barta, Dorina Lakatos, Attila Nagy +2

We introduce HunSum-1: a dataset for Hungarian abstractive summarization, consisting of 1.14M news articles. The dataset is built by collecting, cleaning and deduplicating data fro…

cs.CL2022

Syntax-based data augmentation for Hungarian-English machine translation

Attila Nagy, Patrick Nanys, Balázs Frey Konrád +2

We train Transformer-based neural machine translation models for Hungarian-English and English-Hungarian using the Hunglish2 corpus. Our best models achieve a BLEU score of 40.0 on…