collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2024

Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation for Classification

Jan Cegin, Branislav Pecher, Jakub Simko +3

The generative large language models (LLMs) are increasingly used for data augmentation tasks, where text samples are paraphrased (or generated anew) and then used for classifier f…

cs.CL2024

Multilingual and Multi-topical Benchmark of Fine-tuned Language models and Large Language Models for Check-Worthy Claim Detection

Martin Hyben, Sebastian Kula, Ivan Srba +2

This study compares the performance of (1) fine-tuned language models and (2) large language models on the task of check-worthy claim detection. For the purpose of the comparison w…

cs.CL2024

Authorship Obfuscation in Multilingual Machine-Generated Text Detection

Dominik Macko, Robert Moro, Adaku Uchendu +7

High-quality text generation capability of recent Large Language Models (LLMs) causes concerns about their misuse (e.g., in massive generation/spread of disinformation). Machine-ge…

cs.CL2024

Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation

Branislav Pecher, Jan Cegin, Robert Belanec +3

While fine-tuning of pre-trained language models generally helps to overcome the lack of labelled training samples, it also displays model performance instability. This instability…

cs.CL2024

LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?

Jan Cegin, Jakub Simko, Peter Brusilovsky

The generative large language models (LLMs) are increasingly being used for data augmentation tasks, where text samples are LLM-paraphrased and then used for classifier fine-tuning…

cs.CL2024

Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation

Jan Cegin, Branislav Pecher, Jakub Simko +3

The latest generative large language models (LLMs) have found their application in data augmentation tasks, where small numbers of text samples are LLM-paraphrased and then used to…