Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
How Important is `Perfect' English for Machine Translation Prompts?
Patrícia Schmidtová, Niyati Bafna, Seth Aycock +4
Large language models (LLMs) have achieved top results in recent machine translation evaluations, but they are also known to be sensitive to errors and perturbations in their promp…
cs.CL2025
Conditional Unigram Tokenization with Parallel Data
Gianluca Vico, Jindřinch Libovický
We introduce conditional unigram tokenization, a novel approach that extends unigram tokenization by conditioning target token probabilities on source-language tokens from parallel…
cs.CL2023
Larth: Dataset and Machine Translation for Etruscan
Gianluca Vico, Gerasimos Spanakis
Etruscan is an ancient language spoken in Italy from the 7th century BC to the 1st century AD. There are no native speakers of the language at the present day, and its resources ar…