3 papers
cs.IR2018
Multi-reference Cosine: A New Approach to Text Similarity Measurement in Large Collections
Hamid Mohammadi, Amin Nikoukaran
The importance of an efficient and scalable document similarity detection system is undeniable nowadays. Search engines need batch text similarity measures to detect duplicated and…
cs.CL2018
A Machine Learning Approach to Persian Text Readability Assessment Using a Crowdsourced Dataset
Hamid Mohammadi, Seyed Hossein Khasteh
An automated approach to text readability assessment is essential to a language and can be a powerful tool for improving the understandability of texts written and published in tha…
cs.IR2018
A Fast Text Similarity Measure for Large Document Collections using Multi-reference Cosine and Genetic Algorithm
Hamid Mohammadi, Seyed Hossein Khasteh
One of the important factors that make a search engine fast and accurate is a concise and duplicate free index. In order to remove duplicate and near-duplicate documents from the i…