3 papers
cs.AI2026
Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches
Teddy Ferdinan, BartÅomiej Koptyra, MikoÅaj Langner +42
While Reasoning Language Models (RLMs) are rapidly emerging as powerful tools for scientific research, their impact is primarily concentrated in "hard science" fields. The slow --…
cs.CL2026
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
cs.AI2026
What properties of reasoning supervision are associated with improved downstream model quality?
MikoÅaj Langner, Dzmitry Pihulski, Jan Eliasz +5
Validating training data for reasoning models typically requires expensive trial-and-error fine-tuning cycles. In this work, we investigate whether the utility of a reasoning datas…