3 papers
cs.CL2025
Explaining word embeddings with perfect fidelity: Case study in research impact prediction
Lucie Dvorackova, Marcin P. Joachimiak, Michal Cerny +3
The best-performing approaches for scholarly document quality prediction are based on embedding models. In addition to their performance when used in classifiers, embedding models…
cs.LG2025
LLM-based feature generation from text for interpretable machine learning
VojtÄch Balek, Lukáš Sýkora, Vilém Sklenák +1
Existing text representations such as embeddings and bag-of-words are not suitable for rule learning due to their high dimensionality and absent or questionable feature-level inter…
cs.CL2025
Meaningless is better: hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning
Milena Chadimová, Eduard Jurášek, Tomáš Kliegr
This paper introduces a novel method, referred to as "hashing", which involves masking potentially bias-inducing words in large language models (LLMs) with hash-like meaningless id…