3 papers
math.ST2026
Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model
Insung Kong, Niklas Dexheimer, Johannes Schmidt-Hieber
A refined statistical understanding of LLM pre-training requires the analysis of the transformer architecture for data distributions that encapsulate key characteristics of text da…
cs.LG2026
Hyper Input Convex Neural Networks for Shape Constrained Learning and Optimal Transport
Shayan Hundrieser, Insung Kong, Johannes Schmidt-Hieber
We introduce Hyper Input Convex Neural Networks (HyCNNs), a novel neural network architecture designed for learning convex functions. HyCNNs combine the principles of Maxout networ…
cs.NE2025
On the Universal Representation Property of Spiking Neural Networks
Shayan Hundrieser, Philipp Tuchel, Insung Kong +1
Inspired by biology, spiking neural networks (SNNs) process information via discrete spikes over time, offering an energy-efficient alternative to the classical computing paradigm…