activity
20162026
most citedBiologically plausible deep learning -- but how far can we go with shallow networks?

115 citations · 153 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Teacher Geometry Shapes Learnability in Teacher-Student Networks

Kai J. Sandbrink, Flavio Martinelli, Alexander van Meegen +2

Teacher-student systems, in which a teacher neural network generates training labels so that a student neural network can learn to implement the same function, are widely used as a…

cs.LG2025

Flat Channels to Infinity in Neural Loss Landscapes

Flavio Martinelli, Alexander Van Meegen, Berfin Şimşek +2

The loss landscapes of neural networks contain minima and saddle points that may be connected in flat regions or appear in isolation. We identify and characterize a special structu…

cs.LG2023

Should Under-parameterized Student Networks Copy or Average Teacher Weights?

Berfin Şimşek, Amire Bendjeddou, Wulfram Gerstner +1

Any continuous function can be approximated arbitrarily well by a neural network with sufficiently many neurons . We consider the case when itself is a neural networ…

cs.LG2022

Neural NID Rules

Luca Viano, Johanni Brea

Abstract object properties and their relations are deeply rooted in human common sense, allowing people to predict the dynamics of the world even in situations that are novel but g…

cs.LG2021

Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances

Berfin Şimşek, François Ged, Arthur Jacot +4

We study how permutation symmetries in overparameterized multi-layer neural networks generate `symmetry-induced' critical points. Assuming a network with layers of minimal wi…

cs.LG201917 cited

Weight-space symmetry in deep networks gives rise to permutation saddles, connected by equal-loss valleys across the loss landscape

Johanni Brea, Berfin Simsek, Bernd Illing +1

The permutation symmetry of neurons in each layer of a deep neural network gives rise not only to multiple equivalent global minima of the loss function, but also to first-order sa…