3 papers
math.ST2026
Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model
Insung Kong, Niklas Dexheimer, Johannes Schmidt-Hieber
A refined statistical understanding of LLM pre-training requires the analysis of the transformer architecture for data distributions that encapsulate key characteristics of text da…
cs.LG2026
Spike-timing-dependent Hebbian learning as noisy gradient descent
Niklas Dexheimer, Sascha Gaudlitz, Johannes Schmidt-Hieber
Hebbian learning is a key principle underlying learning in biological neural networks. We relate a Hebbian spike-timing-dependent plasticity rule to noisy gradient descent with res…
math.ST2024
Improving the Convergence Rates of Forward Gradient Descent with Repeated Sampling
Niklas Dexheimer, Johannes Schmidt-Hieber
Forward gradient descent (FGD) has been proposed as a biologically more plausible alternative of gradient descent as it can be computed without backward pass. Considering the linea…