collaborators

6 papers

cs.AI2025

Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine

Wenyi Wang, Piotr Piękos, Li Nanbo +5

Recent studies operationalize self-improvement through coding agents that edit their own codebases. They grow a tree of self-modifications through expansion strategies that favor h…

cs.LG2025

PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors

Yimeng Chen, Piotr Piȩkos, Mateusz Ostaszewski +2

Evaluating the scientific discovery capabilities of large language model based agents, particularly how they cope with varying environmental complexity and utilize prior knowledge,…

cs.LG2025

Is Temporal Difference Learning the Gold Standard for Stitching in RL?

Michał Bortkiewicz, Władysław Pałucki, Mateusz Ostaszewski +1

Reinforcement learning (RL) promises to solve long-horizon tasks even when training data contains only short fragments of the behaviors. This experience stitching capability is oft…

cs.LG2025

Balancing Expressivity and Robustness: Constrained Rational Activations for Reinforcement Learning

Rafał Surdej, Michał Bortkiewicz, Alex Lewandowski +2

Trainable activation functions, whose parameters are optimized alongside network weights, offer increased expressivity compared to fixed activation functions. Specifically, trainab…

cs.LG2025

Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization

Wojciech Masarczyk, Mateusz Ostaszewski, Tin Sum Cheng +3

The softmax function is a fundamental building block of deep neural networks, commonly used to define output distributions in classification tasks or attention weights in transform…

cs.LG2024

Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control

Michal Nauman, Mateusz Ostaszewski, Krzysztof Jankowski +2

Sample efficiency in Reinforcement Learning (RL) has traditionally been driven by algorithmic enhancements. In this work, we demonstrate that scaling can also lead to substantial i…