11 papers
Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model
Insung Kong, Niklas Dexheimer, Johannes Schmidt-Hieber
A refined statistical understanding of LLM pre-training requires the analysis of the transformer architecture for data distributions that encapsulate key characteristics of text da…
Hyper Input Convex Neural Networks for Shape Constrained Learning and Optimal Transport
Shayan Hundrieser, Insung Kong, Johannes Schmidt-Hieber
We introduce Hyper Input Convex Neural Networks (HyCNNs), a novel neural network architecture designed for learning convex functions. HyCNNs combine the principles of Maxout networ…
Central limit theorems for the outputs of fully convolutional neural networks with time series input
Annika Betken, Giorgio Micali, Johannes Schmidt-Hieber
Deep learning is widely deployed for time series learning tasks such as classification and forecasting. Despite the empirical successes, only little theory has been developed so fa…
Semi-Supervised Learning on Graphs using Graph Neural Networks
Juntong Chen, Claire Donnat, Olga Klopp +1
Graph neural networks (GNNs) work remarkably well in semi-supervised node regression, yet a rigorous theory explaining when and why they succeed remains lacking. To address this ga…
Spike-timing-dependent Hebbian learning as noisy gradient descent
Niklas Dexheimer, Sascha Gaudlitz, Johannes Schmidt-Hieber
Hebbian learning is a key principle underlying learning in biological neural networks. We relate a Hebbian spike-timing-dependent plasticity rule to noisy gradient descent with res…
On the Universal Representation Property of Spiking Neural Networks
Shayan Hundrieser, Philipp Tuchel, Insung Kong +1
Inspired by biology, spiking neural networks (SNNs) process information via discrete spikes over time, offering an energy-efficient alternative to the classical computing paradigm…