collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2025

MediQAl: A French Medical Question Answering Dataset for Knowledge and Reasoning Evaluation

Adrien Bazoge

This work introduces MediQAl, a French medical question answering dataset designed to evaluate the capabilities of language models in factual medical recall and reasoning over real…

cs.CL2024

DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain

Yanis Labrak, Adrien Bazoge, Oumaima El Khettari +8

The biomedical domain has sparked a significant interest in the field of Natural Language Processing (NLP), which has seen substantial advancements with pre-trained language models…

cs.CL2024

How Important Is Tokenization in French Medical Masked Language Models?

Yanis Labrak, Adrien Bazoge, Beatrice Daille +2

Subword tokenization has become the prevailing standard in the field of natural language processing (NLP) over recent years, primarily due to the widespread utilization of pre-trai…

cs.CL2024

BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains

Yanis Labrak, Adrien Bazoge, Emmanuel Morin +3

Large Language Models (LLMs) have demonstrated remarkable versatility in recent years, offering potential applications across specialized domains such as healthcare and medicine. D…

cs.CL20234 cited

DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains

Yanis Labrak, Adrien Bazoge, Richard Dufour +4

In recent years, pre-trained language models (PLMs) achieve the best performance on a wide range of natural language processing (NLP) tasks. While the first models were trained on…

cs.CL2023

FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain

Yanis Labrak, Adrien Bazoge, Richard Dufour +4

This paper introduces FrenchMedMCQA, the first publicly available Multiple-Choice Question Answering (MCQA) dataset in French for medical domain. It is composed of 3,105 questions…