5 papers · 1 filter
Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
Giulia Lanzillotta, Felix Sarnthein, Gil Kur +2
The concept of knowledge distillation (KD) describes the training of a student model from a teacher model and is a widely adopted technique in deep learning. However, it is still n…
Understanding and Minimising Outlier Features in Neural Network Training
Bobby He, Lorenzo Noci, Daniele Paliotta +2
Outlier Features (OFs) are neurons whose activation magnitudes significantly exceed the average over a neural network's (NN) width. They are well known to emerge during standard tr…
Hallmarks of Optimization Trajectories in Neural Networks: Directional Exploration and Redundancy
Sidak Pal Singh, Bobby He, Thomas Hofmann +1
We propose a fresh take on understanding the mechanisms of neural networks by analyzing the rich directional structure of optimization trajectories, represented by their pointwise…
Recurrent Distance Filtering for Graph Representation Learning
Yuhui Ding, Antonio Orvieto, Bobby He +1
Graph neural networks based on iterative one-hop message passing have been shown to struggle in harnessing the information from distant nodes effectively. Conversely, graph transfo…
Simplifying Transformer Blocks
Bobby He, Thomas Hofmann
A simple design recipe for deep Transformers is to compose identical building blocks. But standard transformer blocks are far from simple, interweaving attention and MLP sub-blocks…