paper

Disease Identification From Unstructured User Input

arXiv:1905.01987

Abstract

A method to identify probable diseases from the unstructured textual input (eg, health forum posts) by incorporating a lexicographic and semantic feature based two-phase text classification module and a symptom-disease correlation-based similarity measurement module. One notable aspect of my approach was to develop a competent algorithm to extract all inherent features from the data source to make a better decision.

This was an undergraduate research. The hypotheses it proposes is based on a small number of samples and thus, can not be declared significant. To declare it significant, a large number of sample testing is needed. After that, it can be put through