51 citations · 90 across the 12 of their papers we have counts for
18 papers
Exponentially Faster Language Modelling
Peter Belcak, Roger Wattenhofer
Language models only really need to use an exponential fraction of their neurons for individual inferences. As proof, we present UltraFastBERT, a BERT variant that uses 0.3% of its…
SURF: A Generalization Benchmark for GNNs Predicting Fluid Dynamics
Stefan Künzli, Florian Grötschla, Joël Mathys +1
Simulating fluid dynamics is crucial for the design and development process, ranging from simple valves to complex turbomachinery. Accurately solving the underlying physical equati…
Beyond Prompting: Making Pre-trained Language Models Better Zero-shot Learners by Clustering Representations
Yu Fei, Ping Nie, Zhao Meng +2
Recent work has demonstrated that pre-trained language models (PLMs) are zero-shot learners. However, most existing zero-shot methods involve heavy human engineering or complicated…
A Neural Model for Regular Grammar Induction
Peter Belcák, David Hofer, Roger Wattenhofer
Grammatical inference is a classical problem in computational learning theory and a topic of wider influence in natural language processing. We treat grammars as a model of computa…
Periodic Extrapolative Generalisation in Neural Networks
Peter Belcák, Roger Wattenhofer
The learning of the simplest possible computational pattern -- periodicity -- is an open problem in the research of strong generalisation in neural networks. We formalise the probl…
FACT: Learning Governing Abstractions Behind Integer Sequences
Peter Belcák, Ard Kastrati, Flavio Schenker +1
Integer sequences are of central importance to the modeling of concepts admitting complete finitary descriptions. We introduce a novel view on the learning of such concepts and lay…