BioRED: A Rich Biomedical Relation Extraction Dataset
arXiv:2204.04263 · doi:10.1093/bib/bbac282
Abstract
Automated relation extraction (RE) from biomedical literature is critical for many downstream text mining applications in both research and real-world settings. However, most existing benchmarking datasets for bio-medical RE only focus on relations of a single type (e.g., protein-protein interactions) at the sentence level, greatly limiting the development of RE systems in biomedicine. In this work, we first review commonly used named entity recognition (NER) and RE datasets. Then we present BioRED, a first-of-its-kind biomedical RE corpus with multiple entity types (e.g., gene/protein, disease, chemical) and relation pairs (e.g., gene-disease; chemical-chemical) at the document level, on a set of 600 PubMed abstracts. Further, we label each relation as describing either a novel finding or previously known background knowledge, enabling automated algorithms to differentiate between novel and background information. We assess the utility of BioRED by benchmarking several existing state-of-the-art methods, including BERT-based models, on the NER and RE tasks. Our results show that while existing approaches can reach high performance on the NER task (F-score of 89.3%), there is much room for improvement for the RE task, especially when extracting novel relations (F-score of 47.7%). Our experiments also demonstrate that such a rich dataset can successfully facilitate the development of more accurate, efficient, and robust RE systems for biomedicine. The BioRED dataset and annotation guideline are freely available at https://ftp.ncbi.nlm.nih.gov/pub/lu/BioRED/.
Accepted by Briefings in Bioinformatics
References in corpus (1)
Cited by in corpus (14)
- Taiyi: A Bilingual Fine-Tuned Large Language Model for Diverse Biomedical Tasks
- AIONER: All-in-one scheme-based biomedical named entity recognition using deep learning
- Advancing Chinese biomedical text mining with community challenges
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- Localising In-Domain Adaptation of Transformer-Based Biomedical Language Models
- From Zero to Hero: Harnessing Transformers for Biomedical Named Entity Recognition in Zero- and Few-shot Contexts
- HunFlair2 in a cross-corpus evaluation of biomedical named entity recognition and normalization tools
- Traceable LLM-based validation of statements in knowledge graphs
- A survey on cutting-edge relation extraction techniques based on language models
- DiMB-RE: Mining the Scientific Literature for Diet-Microbiome Associations
- Augmenting Biomedical Named Entity Recognition with General-domain Resources
- Enhancing Biomedical Knowledge Discovery for Diseases: An Open-Source Framework Applied on Rett Syndrome and Alzheimer's Disease
- Data-Driven Information Extraction and Enrichment of Molecular Profiling Data for Cancer Cell Lines
- Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study