3 papers
cs.CL2025
SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
Muskaan Chopra, Lorenz Sparrenberg, Rafet Sifa
Critical Error Detection (CED) in machine translation aims to determine whether a translation is safe to use or contains unacceptable deviations in meaning. While the WMT21 English…
cs.CL2024
Pointer-Guided Pre-Training: Infusing Large Language Models with Paragraph-Level Contextual Awareness
Lars Hillebrand, Prabhupad Pradhan, Christian Bauckhage +1
We introduce "pointer-guided segment ordering" (SO), a novel pre-training technique aimed at enhancing the contextual understanding of paragraph-level text representations in large…
cs.CL2023
Controlled Randomness Improves the Performance of Transformer Models
Tobias Deußer, Cong Zhao, Wolfgang Krämer +3
During the pre-training step of natural language models, the main objective is to learn a general representation of the pre-training dataset, usually requiring large amounts of tex…