7 papers
Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval
You Zuo, Kim Gerdes, Éric de la Clergerie +1
Patent prior-art retrieval is a recall-oriented search task over long and highly structured technical documents. Dense retrieval improves semantic matching, but single-vector repre…
OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining
Rian Touchent, Ãric de la Clergerie
We present OntoBook, a method that converts medical ontology structure into pretraining signal for encoder language models. Our approach has three stages: random walks through onto…
Patent Representation Learning via Self-supervision
You Zuo, Kim Gerdes, Eric Villemonte de La Clergerie +1
We study self-supervised patent representation learning with contrastive objectives. A standard baseline constructs positives by encoding the same text twice under independent drop…
A Causal Language Modeling Detour Improves Encoder Continued Pretraining
Rian Touchent, Eric de la Clergerie
When adapting an encoder to a new domain, the standard approach is to continue training with Masked Language Modeling (MLM). We show that temporarily switching to Causal Language M…
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
Rian Touchent, Nathan Godey, Eric de la Clergerie
We introduce Biomed-Enriched, a biomedical text dataset constructed from PubMed via a two-stage annotation process. In the first stage, a large language model annotates 400K paragr…
PatentEval: Understanding Errors in Patent Generation
You Zuo, Kim Gerdes, Eric Villemonte de La Clergerie +1
In this work, we introduce a comprehensive error typology specifically designed for evaluating two distinct tasks in machine-generated patent texts: claims-to-abstract generation,…