4 papers
Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model
Insung Kong, Niklas Dexheimer, Johannes Schmidt-Hieber
A refined statistical understanding of LLM pre-training requires the analysis of the transformer architecture for data distributions that encapsulate key characteristics of text da…
Hyper Input Convex Neural Networks for Shape Constrained Learning and Optimal Transport
Shayan Hundrieser, Insung Kong, Johannes Schmidt-Hieber
We introduce Hyper Input Convex Neural Networks (HyCNNs), a novel neural network architecture designed for learning convex functions. HyCNNs combine the principles of Maxout networ…
On the Universal Representation Property of Spiking Neural Networks
Shayan Hundrieser, Philipp Tuchel, Insung Kong +1
Inspired by biology, spiking neural networks (SNNs) process information via discrete spikes over time, offering an energy-efficient alternative to the classical computing paradigm…
On the expressivity of deep Heaviside networks
Insung Kong, Juntong Chen, Sophie Langer +1
We show that deep Heaviside networks (DHNs) have limited expressiveness but that this can be overcome by including either skip connections or neurons with linear activation. We pro…