5 papers
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
Mihai Nadas, Laura Diosan, Andrei Piscoran +1
Moral stories are a time-tested vehicle for transmitting values, yet modern NLP lacks a large, structured corpus that couples coherent narratives with explicit ethical lessons. We…
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
Mihai Nadas, Laura Diosan, Andreea Tomescu +1
Literary translation has recently gained attention as a distinct and complex task in machine translation research, yet translation by small open models remains an open problem, par…
TF3-RO-50M: Training Compact Romanian Language Models from Scratch on Synthetic Moral Microfiction
Mihai Dan Nadas, Laura Diosan, Andreea Tomescu +1
Recent advances in synthetic data generation have shown that compact language models can be trained effectively when the underlying corpus is structurally controlled and linguistic…
Evaluating Large Language Models for Diacritic Restoration in Romanian Texts: A Comparative Study
Mihai Nadas, Laura Diosan
Automatic diacritic restoration is crucial for text processing in languages with rich diacritical marks, such as Romanian. This study evaluates the performance of several large lan…
Synthetic Data Generation Using Large Language Models: Advances in Text and Code
Mihai Nadas, Laura Diosan, Andreea Tomescu
This survey reviews how large language models (LLMs) are transforming synthetic training data generation in both natural language and code domains. By producing artificial but task…