papers

Publications (8)

cs.CL2021

Clustering-based Inference for Biomedical Entity Linking

Rico Angell, Nicholas Monath, Sunil Mohan +2

Due to large number of entities in biomedical knowledge bases, only a small fraction of entities have corresponding labelled training data. This necessitates entity linking models…

cs.CL2019

MedMentions: A Large Biomedical Corpus Annotated with UMLS Concepts

Sunil Mohan, Donghui Li

This paper presents the formal release of MedMentions, a new manually annotated resource for the recognition of biomedical concepts. What distinguishes MedMentions from other annot…

cs.CL2022

A Distant Supervision Corpus for Extracting Biomedical Relationships Between Chemicals, Diseases and Genes

Dongxu Zhang, Sunil Mohan, Michaela Torkar +1

We introduce ChemDisGene, a new dataset for training and evaluating multi-class multi-label document-level biomedical relation extraction models. Our dataset contains 80k biomedica…

cs.CL2021

Low Resource Recognition and Linking of Biomedical Concepts from a Large Ontology

Sunil Mohan, Rico Angell, Nick Monath +1

Tools to explore scientific literature are essential for scientists, especially in biomedicine, where about a million new papers are published every year. Many such tools provide u…

cs.CL2022

LSTM-RASA Based Agri Farm Assistant for Farmers

Narayana Darapaneni, Selvakumar Raj, Raghul V +3

The application of Deep Learning and Natural Language based ChatBots are growing rapidly in recent years. They are used in many fields like customer support, reservation system and…

cs.CL2020

Using Error Decay Prediction to Overcome Practical Issues of Deep Active Learning for Named Entity Recognition

Haw-Shiuan Chang, Shankar Vembu, Sunil Mohan +2

Existing deep active learning algorithms achieve impressive sampling efficiency on natural language processing tasks. However, they exhibit several weaknesses in practice, includin…

cs.CL2025

How Well Do LLMs Understand Drug Mechanisms? A Knowledge + Reasoning Evaluation Dataset

Sunil Mohan, Theofanis Karaletsos

Two scientific fields showing increasing interest in pre-trained large language models (LLMs) are drug development / repurposing, and personalized medicine. For both, LLMs have to…

cs.IR2018

A Fast Deep Learning Model for Textual Relevance in Biomedical Information Retrieval

Sunil Mohan, Nicolas Fiorini, Sun Kim +1

Publications in the life sciences are characterized by a large technical vocabulary, with many lexical and semantic variations for expressing the same concept. Towards addressing t…