24 citations · 25 across the 5 of their papers we have counts for
6 papers · 1 filter
From Words to Amino Acids: Does the Curse of Depth Persist?
Aleena Siji, Amir Mohammad Karimi Mamaghan, Ferdinand Kapl +9
Protein language models (PLMs) have become widely adopted as general-purpose models, demonstrating strong performance in protein engineering and de novo design. Like large language…
Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
Seijin Kobayashi, Yanick Schimpf, Maximilian Schlegel +12
Large-scale autoregressive models pretrained on next-token prediction and finetuned with reinforcement learning (RL) have achieved unprecedented success on many problem domains. Du…
Learning Randomized Algorithms with Transformers
Johannes von Oswald, Seijin Kobayashi, Yassir Akram +1
Randomization is a powerful tool that endows algorithms with remarkable properties. For instance, randomized algorithms excel in adversarial settings, often surpassing the worst-ca…
Disentangling the Predictive Variance of Deep Ensembles through the Neural Tangent Kernel
Seijin Kobayashi, Pau Vilimelis Aceituno, Johannes von Oswald
Identifying unfamiliar inputs, also known as out-of-distribution (OOD) detection, is a crucial property of any decision making process. A simple and empirically validated technique…
Learning where to learn: Gradient sparsity in meta and continual learning
Johannes von Oswald, Dominic Zhao, Seijin Kobayashi +4
Finding neural network weights that generalize well from small datasets is difficult. A promising approach is to learn a weight initialization such that a small number of weight ch…
Continual Learning in Recurrent Neural Networks
Benjamin Ehret, Christian Henning, Maria R. Cervera +3
While a diverse collection of continual learning (CL) methods has been proposed to prevent catastrophic forgetting, a thorough investigation of their effectiveness for processing s…