collaborators

6 papers

cs.CL2025

Scaling Data-Constrained Language Models

Niklas Muennighoff, Alexander M. Rush, Boaz Barak +6

The current trend of scaling language models involves increasing both parameter count and training dataset size. Extrapolating this trend suggests that training dataset size may so…

cs.CL2025

Poro 34B and the Blessing of Multilinguality

Risto Luukkonen, Jonathan Burdge, Elaine Zosa +5

The pretraining of state-of-the-art large language models now requires trillions of words of text, which is orders of magnitude more than available for the vast majority of languag…

cs.CL2025

Got Compute, but No Data: Lessons From Post-training a Finnish LLM

Elaine Zosa, Ville Komulainen, Sampo Pyysalo

As LLMs gain more popularity as chatbots and general assistants, methods have been developed to enable LLMs to follow instructions and align with human preferences. These methods h…

cs.CL2024

Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code

Taishi Nakamura, Mayank Mishra, Simone Tedeschi +42

Pretrained language models are an integral part of AI applications, but their high computational cost for training limits accessibility. Initiatives such as Bloom and StarCoder aim…

cond-mat.mtrl-sci2024

Question Answering models for information extraction from perovskite materials science literature

M. Sipilä, F. Mehryary, S. Pyysalo +2

Scientific text is a promising source of data in materials science, with ongoing research into utilising textual data for materials discovery. In this study, we developed and teste…

cs.CL2024

A Survey of Large Language Models for European Languages

Wazir Ali, Sampo Pyysalo

Large Language Models (LLMs) have gained significant attention due to their high performance on a wide range of natural language tasks since the release of ChatGPT. The LLMs learn…