6 papers
Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
Wenyi Wang, Piotr PiÄkos, Li Nanbo +5
Recent studies operationalize self-improvement through coding agents that edit their own codebases. They grow a tree of self-modifications through expansion strategies that favor h…
PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
Yimeng Chen, Piotr Piȩkos, Mateusz Ostaszewski +2
Evaluating the scientific discovery capabilities of large language model based agents, particularly how they cope with varying environmental complexity and utilize prior knowledge,…
Is Temporal Difference Learning the Gold Standard for Stitching in RL?
MichaÅ Bortkiewicz, WÅadysÅaw PaÅucki, Mateusz Ostaszewski +1
Reinforcement learning (RL) promises to solve long-horizon tasks even when training data contains only short fragments of the behaviors. This experience stitching capability is oft…
Balancing Expressivity and Robustness: Constrained Rational Activations for Reinforcement Learning
RafaÅ Surdej, MichaÅ Bortkiewicz, Alex Lewandowski +2
Trainable activation functions, whose parameters are optimized alongside network weights, offer increased expressivity compared to fixed activation functions. Specifically, trainab…
Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization
Wojciech Masarczyk, Mateusz Ostaszewski, Tin Sum Cheng +3
The softmax function is a fundamental building block of deep neural networks, commonly used to define output distributions in classification tasks or attention weights in transform…
Bigger, Regularized, Optimistic: scaling for compute and sample-efficient continuous control
Michal Nauman, Mateusz Ostaszewski, Krzysztof Jankowski +2
Sample efficiency in Reinforcement Learning (RL) has traditionally been driven by algorithmic enhancements. In this work, we demonstrate that scaling can also lead to substantial i…