4 papers
Text classification with word embedding regularization and soft similarity measure
Vít Novotný, Eniafe Festus Ayetiran, Michal Štefánik +1
Since the seminal work of Mikolov et al., word embeddings have become the preferred word representations for many natural language processing tasks. Document similarity measures ex…
Implementation Notes for the Soft Cosine Measure
Vít Novotný
The standard bag-of-words vector space model (VSM) is efficient, and ubiquitous in information retrieval, but it underestimates the similarity of documents with the same meaning, b…
MIaS: Math-Aware Retrieval in Digital Mathematical Libraries
Petr Sojka, Michal Růžička, Vít Novotný
Digital mathematical libraries (DMLs) such as arXiv, Numdam, and EuDML contain mainly documents from STEM fields, where mathematical formulae are often more important than text for…
Semantic Vector Encoding and Similarity Search Using Fulltext Search Engines
Jan Rygl, Jan Pomikálek, Radim Řehůřek +3
Vector representations and vector space modeling (VSM) play a central role in modern machine learning. We propose a novel approach to `vector similarity searching' over dense seman…