3 papers
cs.CL2026
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
Mihai Nadas, Laura Diosan, Andrei Piscoran +1
Moral stories are a time-tested vehicle for transmitting values, yet modern NLP lacks a large, structured corpus that couples coherent narratives with explicit ethical lessons. We…
cs.CL2026
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
Mihai Nadas, Laura Diosan, Andreea Tomescu +1
Literary translation has recently gained attention as a distinct and complex task in machine translation research, yet translation by small open models remains an open problem, par…
cs.CL2026
TF3-RO-50M: Training Compact Romanian Language Models from Scratch on Synthetic Moral Microfiction
Mihai Dan Nadas, Laura Diosan, Andreea Tomescu +1
Recent advances in synthetic data generation have shown that compact language models can be trained effectively when the underlying corpus is structurally controlled and linguistic…