15 citations · 15 across the 13 of their papers we have counts for
14 papers · 1 filter
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina +2
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced d…
Modeling semantic association in self-paced reading with language model embeddings
Sara Møller Ãstergaard, Kenneth Enevoldsen, Afra Alishahi +1
Semantic association between a word and its context has been identified as an important component of reading comprehension, even when word predictability is accounted for. Recent r…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
Grounding Text Embeddings in Stakeholder Associations
Jonathan Rystrøm, Sofie Burgos-Thorsen, Zihao Fu +3
Text embeddings are widely used to analyse large corpora of complex texts. However, it is unclear whether the embeddings capture the same semantic distances as the human experts us…
Naturalistic measure of social norms alignment
Yevhen Kostiuk, Kenneth Enevoldsen, Peter Bjerregaard Vahlstrup +2
Social norms reflect shared expectations on acceptable behavior. Measuring social norms alignment remains challenging, with existing approaches typically relying on artificial clos…
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
Yevhen Kostiuk, Kenneth Enevoldsen
Instruction embedding models have become common among state-of-the-art models, however are evaluated using a single prompt per task. The single-point evaluation ignores a main prob…