3 papers
cs.CL2025
The PLLuM Instruction Corpus
Piotr PÄzik, Filip Å»arnecki, Konrad KaczyÅski +50
This paper describes the instruction dataset used to fine-tune a set of transformer-based large language models (LLMs) developed in the PLLuM (Polish Large Language Model) project.…
cs.CL2025
LLMLagBench: Identifying Temporal Training Boundaries in Large Language Models
Piotr PÄzik, Konrad KaczyÅski, Maria SzymaÅska +4
Large Language Models (LLMs) are pretrained on textual data up to a specific temporal cutoff. This creates a strict knowledge boundary beyond which models cannot provide accurate i…
cs.CL2025
PLLuM: A Family of Polish Large Language Models
Jan KocoÅ, Maciej Piasecki, Arkadiusz Janz +96
Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for ot…