activity
20242026
collaborators

11 papers

cs.LG2026

Timestep-Conditioned Transformers for Global Weather Forecasting

Sam Levang, Fran Bartolic, Ty Dickinson +3

Existing machine-learning weather forecasting models rely on predetermined and fixed autoregressive timesteps. The choice of model timestep involves a fundamental trade-off: shorte…

cs.AI2026

Demonstrating Generalization Failures via Mixtures of Conditional Policies

Jou Barzdukas, Jack Peck, Julian Schulz +3

Post-training of frontier language models is conducted on curated task suites, and inevitably leaves a distribution shift between training and deployment environments. This exposes…

cs.LG2026

Recursive Scaling in Masked Diffusion Models

Alba Carballo-Castro, Julianna Piskorz, Paulius Rauba +2

Masked diffusion models (MDMs) have recently emerged as a promising paradigm for sequence generation. Scaling MDMs is conventionally achieved by increasing the parameter count or t…

cs.LG2026

Tiny Autoregressive Recursive Models

Paulius Rauba, Claudio Fanconi, Mihaela van der Schaar

Tiny Recursive Models (TRMs) have recently demonstrated remarkable performance on ARC-AGI, showing that very small models can compete against large foundation models through a two-…

cs.CR2026

No More, No Less: Least-Privilege Language Models

Paulius Rauba, Dominykas Seputis, Patrikas Vanagas +1

Least privilege is a core security principle: grant each request only the minimum access needed to achieve its goal. Deployed language models almost never follow it, instead being…

cs.LG2026

Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models

Paulius Rauba, Mihaela van der Schaar

Large neural networks are typically trained for a fixed computational budget, creating a rigid trade-off between performance and efficiency that is ill-suited for deployment in res…