11 papers
Timestep-Conditioned Transformers for Global Weather Forecasting
Sam Levang, Fran Bartolic, Ty Dickinson +3
Existing machine-learning weather forecasting models rely on predetermined and fixed autoregressive timesteps. The choice of model timestep involves a fundamental trade-off: shorte…
Demonstrating Generalization Failures via Mixtures of Conditional Policies
Jou Barzdukas, Jack Peck, Julian Schulz +3
Post-training of frontier language models is conducted on curated task suites, and inevitably leaves a distribution shift between training and deployment environments. This exposes…
Recursive Scaling in Masked Diffusion Models
Alba Carballo-Castro, Julianna Piskorz, Paulius Rauba +2
Masked diffusion models (MDMs) have recently emerged as a promising paradigm for sequence generation. Scaling MDMs is conventionally achieved by increasing the parameter count or t…
Tiny Autoregressive Recursive Models
Paulius Rauba, Claudio Fanconi, Mihaela van der Schaar
Tiny Recursive Models (TRMs) have recently demonstrated remarkable performance on ARC-AGI, showing that very small models can compete against large foundation models through a two-…
No More, No Less: Least-Privilege Language Models
Paulius Rauba, Dominykas Seputis, Patrikas Vanagas +1
Least privilege is a core security principle: grant each request only the minimum access needed to achieve its goal. Deployed language models almost never follow it, instead being…
Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
Paulius Rauba, Mihaela van der Schaar
Large neural networks are typically trained for a fixed computational budget, creating a rigid trade-off between performance and efficiency that is ill-suited for deployment in res…