Identification of Protein Coding Regions in Genomic DNA Using Unsupervised FMACA Based Pattern Classifier
arXiv:1401.6484
Abstract
Genes carry the instructions for making proteins that are found in a cell as a specific sequence of nucleotides that are found in DNA molecules. But, the regions of these genes that code for proteins may occupy only a small region of the sequence. Identifying the coding regions play a vital role in understanding these genes. In this paper we propose a unsupervised Fuzzy Multiple Attractor Cellular Automata (FMCA) based pattern classifier to identify the coding region of a DNA sequence. We propose a distinct K-Means algorithm for designing FMACA classifier which is simple, efficient and produces more accurate classifier than that has previously been obtained for a range of different sequence lengths. Experimental results confirm the scalability of the proposed Unsupervised FCA based classifier to handle large volume of datasets irrespective of the number of classes, tuples and attributes. Good classification accuracy has been established.
arXiv admin note: text overlap with arXiv:1312.2642
Cited by in corpus (5)
- Cellular Automata and Its Applications in Bioinformatics: A Review
- Multiple Attractor Cellular Automata (MACA) for Addressing Major Problems in Bioinformatics
- An Extensive Report on Cellular Automata Based Artificial Immune System for Strengthening Automated Protein Prediction
- Cellular Automata based Feedback Mechanism in Strengthening biological Sequence Analysis Approach to Robotic Soccer
- CAVDM: Cellular Automata Based Video Cloud Mining Framework for Information Retrieval