6 papers
Naturalistic measure of social norms alignment
Yevhen Kostiuk, Kenneth Enevoldsen, Peter Bjerregaard Vahlstrup +2
Social norms reflect shared expectations on acceptable behavior. Measuring social norms alignment remains challenging, with existing approaches typically relying on artificial clos…
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
Yevhen Kostiuk, Kenneth Enevoldsen
Instruction embedding models have become common among state-of-the-art models, however are evaluated using a single prompt per task. The single-point evaluation ignores a main prob…
Belief Propagation in LLM World Models: Measuring Strategic Information Bias with Prediction Markets
Mykola Khandoga, Yevhen Kostiuk, Anton Polishko +4
Every information ecosystem produces beliefs that shape strategic decisions. Both human analysts and AI systems inherit the blind spots of their information sources. We show that L…
The Veln(ia)s is in the Details: Evaluating LLM Judgment on Latvian and Lithuanian Short Answer Matching
Yevhen Kostiuk, Oxana Vitman, Åukasz GagaÅa +1
In this work, we address the challenge of evaluating large language models (LLMs) on the short answer matching task for Latvian and Lithuanian languages. We introduce novel dataset…
Towards Multilingual LLM Evaluation for Baltic and Nordic languages: A study on Lithuanian History
Yevhen Kostiuk, Oxana Vitman, Åukasz GagaÅa +1
In this work, we evaluated Lithuanian and general history knowledge of multilingual Large Language Models (LLMs) on a multiple-choice question-answering task. The models were teste…
From English-Centric to Effective Bilingual: LLMs with Custom Tokenizers for Underrepresented Languages
Artur Kiulian, Anton Polishko, Mykola Khandoga +10
In this paper, we propose a model-agnostic cost-effective approach to developing bilingual base large language models (LLMs) to support English and any target language. The method…