3 papers
cs.CL2025
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
Joona Kytöniemi, Jousia Piha, Akseli Reunamo +3
We introduce FIN-bench-v2, a unified benchmark suite for evaluating large language models in Finnish. FIN-bench-v2 consolidates Finnish versions of widely used benchmarks together…
cs.CL2025
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
Laurie Burchell, Ona de Gibert, Nikolay Arefyev +32
Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In th…
cond-mat.mtrl-sci2024
Question Answering models for information extraction from perovskite materials science literature
M. Sipilä, F. Mehryary, S. Pyysalo +2
Scientific text is a promising source of data in materials science, with ongoing research into utilising textual data for materials discovery. In this study, we developed and teste…