activity
20202022
most citedAutomatic punctuation restoration with BERT models

8 citations · 11 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2022

Syntax-based data augmentation for Hungarian-English machine translation

Attila Nagy, Patrick Nanys, Balázs Frey Konrád +2

We train Transformer-based neural machine translation models for Hungarian-English and English-Hungarian using the Hunglish2 corpus. Our best models achieve a BLEU score of 40.0 on…

cs.CL2021

A Three Step Training Approach with Data Augmentation for Morphological Inflection

Gabor Szolnok, Botond Barta, Dorina Lakatos +1

We present the BME submission for the SIGMORPHON 2021 Task 0 Part 1, Generalization Across Typologically Diverse Languages shared task. We use an LSTM encoder-decoder model with th…

cs.CL2021

Subword Pooling Makes a Difference

Judit Ács, Ákos Kádár, András Kornai

Contextual word-representations became a standard in modern natural language processing systems. These models use subword tokenization to handle large vocabularies and unknown word…

cs.CL20213 cited

Evaluating Contextualized Language Models for Hungarian

Judit Ács, Dániel Lévai, Dávid Márk Nemeskey +1

We present an extended comparison of contextualized language models for Hungarian. We compare huBERT, a Hungarian model against 4 multilingual models including the multilingual BER…

cs.CL20218 cited

Automatic punctuation restoration with BERT models

Attila Nagy, Bence Bial, Judit Ács

We present an approach for automatic punctuation restoration with BERT models for English and Hungarian. For English, we conduct our experiments on Ted Talks, a commonly used bench…

cs.CL2020

The Role of Interpretable Patterns in Deep Learning for Morphology

Judit Acs, Andras Kornai

We examine the role of character patterns in three tasks: morphological analysis, lemmatization and copy. We use a modified version of the standard sequence-to-sequence model, wher…