activity
20242026
collaborators

7 papers

cs.CL2026

Latent Knowledge as a Predictor of Fact Acquisition in Fine-Tuned Large Language Models

Daniel B. Hier, Tayo Obafemi-Ajayi

Large language models store biomedical facts with uneven strength after pretraining: some facts are present in the weights but are not reliably accessible under deterministic decod…

cs.CL2026

Predicting Failures of LLMs to Link Biomedical Ontology Terms to Identifiers Evidence Across Models and Ontologies

Daniel B. Hier, Steven Keith Platt, Tayo Obafemi-Ajayi

Large language models often perform well on biomedical NLP tasks but may fail to link ontology terms to their correct identifiers. We investigate why these failures occur by analyz…

cs.CL2025

From Memorization to Generalization: Fine-Tuning Large Language Models for Biomedical Term-to-Identifier Normalization

Suswitha Pericharla, Daniel B. Hier, Tayo Obafemi-Ajayi

Effective biomedical data integration depends on automated term normalization, the mapping of natural language biomedical terms to standardized identifiers. This linking of terms t…

cs.CL2025

High-Throughput Phenotyping of Clinical Text Using Large Language Models

Daniel B. Hier, S. Ilyas Munzir, Anne Stahlfeld +2

High-throughput phenotyping automates the mapping of patient signs to standardized ontology concepts and is essential for precision medicine. This study evaluates the automation of…

cs.CL2025

Mapping Biomedical Ontology Terms to IDs: Effect of Domain Prevalence on Prediction Accuracy

Thanh Son Do, Daniel B. Hier, Tayo Obafemi-Ajayi

This study evaluates the ability of large language models (LLMs) to map biomedical ontology terms to their corresponding ontology IDs across the Human Phenotype Ontology (HPO), Gen…

cs.CL2025

A Simplified Retriever to Improve Accuracy of Phenotype Normalizations by Large Language Models

Daniel B. Hier, Thanh Son Do, Tayo Obafemi-Ajayi

Large language models (LLMs) have shown improved accuracy in phenotype term normalization tasks when augmented with retrievers that suggest candidate normalizations based on term d…