15 papers
Left properness of Moore flows
Philippe Gaucher
We introduce the notion of a reparametrization category with cuts. For every such reparametrization category , we prove the tensor lemma, namely that the tensor product…
Automated Background Swapping for Robustness against Spurious Backgrounds
Cesar Roder, Kajetan Schweighofer
Classifiers based on Deep Neural Networks exhibit strong performance across domains, yet can fail catastrophically if they rely on spurious correlations, i.e., features that are pr…
RREDCoT: Segment-Level Reward Redistribution for Reasoning Models
Mykyta Ielanskyi, Kajetan Schweighofer, Lukas Aichberger +1
Recent advancements in reasoning language models have been driven by Reinforcement Learning (RL) fine-tuning. Most often, these rely on the Group Relative Policy Optimization (GRPO…
Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies
Kajetan Schweighofer, Conor F. Hayes, Roberto Dailey +2
Evolution Strategies (ES) has recently emerged as a competitive alternative to reinforcement learning (RL) for large language model (LLM) fine-tuning, offering advantages through s…
Efficient Pre-Training of LLMs through Truncated SVD Layers
Kaivan Kamali, Kajetan Schweighofer, Hormoz Shahrzad +3
The massive scaling of Large Language Models (LLMs) has made pretraining increasingly cost-prohibitive. While low-rank representation and orthonormal weight matrices could in princ…
Position: agentic AI orchestration should be Bayes-consistent
Theodore Papamarkou, Pierre Alquier, Matthias Bauer +27
LLMs excel at predictive tasks and complex reasoning tasks, but many high-value deployments rely on decisions under uncertainty, for example, which tool to call, which expert to co…