16 papers
Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval
You Zuo, Kim Gerdes, Éric de la Clergerie +1
Patent prior-art retrieval is a recall-oriented search task over long and highly structured technical documents. Dense retrieval improves semantic matching, but single-vector repre…
Patent Representation Learning via Self-supervision
You Zuo, Kim Gerdes, Eric Villemonte de La Clergerie +1
We study self-supervised patent representation learning with contrastive objectives. A standard baseline constructs positives by encoding the same text twice under independent drop…
Translation Heads: Disentangling meaning from language in LLM-based machine translation
Théo Lasnier, Armel Zebaze, Djamé Seddah +2
Mechanistic Interpretability (MI) seeks to explain how neural networks implement their capabilities, but the scale of Large Language Models (LLMs) has limited prior MI work in Mach…
When the Gold Standard Isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content
Lydia Nishimwe, Benoît Sagot, Rachel Bawden
User-generated content (UGC) is characterised by frequent use of non-standard language, from spelling errors to expressive choices such as slang, character repetitions, and emojis.…
Testing the Deliteralization Hypothesis in Human and Machine Translation
Malik Marmonier, Rachel Bawden, Benoît Sagot
The recent shift from dedicated NMT systems to general-purpose LLMs has reshaped machine translation, with LLMs reported to produce more fluent, less literal output than their pred…
Hindsight Quality Prediction Experiments in Multi-Candidate Human-Post-Edited Machine Translation
Malik Marmonier, Benoît Sagot, Rachel Bawden
This paper investigates two complementary paradigms for predicting machine translation (MT) quality: source-side difficulty prediction and candidate-side quality estimation (QE). T…