activity
20242026
collaborators

7 papers

cs.IR2026

Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval

You Zuo, Kim Gerdes, Éric de la Clergerie +1

Patent prior-art retrieval is a recall-oriented search task over long and highly structured technical documents. Dense retrieval improves semantic matching, but single-vector repre…

cs.AI2026

OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining

Rian Touchent, Éric de la Clergerie

We present OntoBook, a method that converts medical ontology structure into pretraining signal for encoder language models. Our approach has three stages: random walks through onto…

cs.CL2026

Patent Representation Learning via Self-supervision

You Zuo, Kim Gerdes, Eric Villemonte de La Clergerie +1

We study self-supervised patent representation learning with contrastive objectives. A standard baseline constructs positives by encoding the same text twice under independent drop…

cs.CL2026

A Causal Language Modeling Detour Improves Encoder Continued Pretraining

Rian Touchent, Eric de la Clergerie

When adapting an encoder to a new domain, the standard approach is to continue training with Masked Language Modeling (MLM). We show that temporarily switching to Causal Language M…

cs.CL2025

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content

Rian Touchent, Nathan Godey, Eric de la Clergerie

We introduce Biomed-Enriched, a biomedical text dataset constructed from PubMed via a two-stage annotation process. In the first stage, a large language model annotates 400K paragr…

cs.CL2024

PatentEval: Understanding Errors in Patent Generation

You Zuo, Kim Gerdes, Eric Villemonte de La Clergerie +1

In this work, we introduce a comprehensive error typology specifically designed for evaluating two distinct tasks in machine-generated patent texts: claims-to-abstract generation,…