collaborators

11 papers

math.ST2026

Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model

Insung Kong, Niklas Dexheimer, Johannes Schmidt-Hieber

A refined statistical understanding of LLM pre-training requires the analysis of the transformer architecture for data distributions that encapsulate key characteristics of text da…

cs.LG2026

Hyper Input Convex Neural Networks for Shape Constrained Learning and Optimal Transport

Shayan Hundrieser, Insung Kong, Johannes Schmidt-Hieber

We introduce Hyper Input Convex Neural Networks (HyCNNs), a novel neural network architecture designed for learning convex functions. HyCNNs combine the principles of Maxout networ…

stat.ME2026

Central limit theorems for the outputs of fully convolutional neural networks with time series input

Annika Betken, Giorgio Micali, Johannes Schmidt-Hieber

Deep learning is widely deployed for time series learning tasks such as classification and forecasting. Despite the empirical successes, only little theory has been developed so fa…

stat.ML2026

Semi-Supervised Learning on Graphs using Graph Neural Networks

Juntong Chen, Claire Donnat, Olga Klopp +1

Graph neural networks (GNNs) work remarkably well in semi-supervised node regression, yet a rigorous theory explaining when and why they succeed remains lacking. To address this ga…

cs.LG2026

Spike-timing-dependent Hebbian learning as noisy gradient descent

Niklas Dexheimer, Sascha Gaudlitz, Johannes Schmidt-Hieber

Hebbian learning is a key principle underlying learning in biological neural networks. We relate a Hebbian spike-timing-dependent plasticity rule to noisy gradient descent with res…

cs.NE2025

On the Universal Representation Property of Spiking Neural Networks

Shayan Hundrieser, Philipp Tuchel, Insung Kong +1

Inspired by biology, spiking neural networks (SNNs) process information via discrete spikes over time, offering an energy-efficient alternative to the classical computing paradigm…