13 citations · 32 across the 4 of their papers we have counts for
5 papers
On the geometry of generalization and memorization in deep neural networks
Cory Stephenson, Suchismita Padhy, Abhinav Ganesh +3
Understanding how large neural networks avoid memorizing training data is key to explaining their high generalization performance. To examine the structure of when and where memori…
Syntactic Perturbations Reveal Representational Correlates of Hierarchical Phrase Structure in Pretrained Language Models
Matteo Alleman, Jonathan Mamou, Miguel A Del Rio +3
While vector-based language representations from pretrained language models have set a new standard for many NLP tasks, there is not yet a complete accounting of their inner workin…
1-bit LAMB: Communication Efficient Large-Scale Large-Batch Training with LAMB's Convergence Speed
Conglong Li, Ammar Ahmad Awan, Hanlin Tang +2
To train large models (like BERT and GPT-3) on hundreds of GPUs, communication has become a major bottleneck, especially on commodity systems with limited-bandwidth TCP network. On…
Emergence of Separable Manifolds in Deep Language Representations
Jonathan Mamou, Hang Le, Miguel Del Rio +4
Deep neural networks (DNNs) have shown much empirical success in solving perceptual tasks across various cognitive modalities. While they are only loosely inspired by the biologica…
Untangling in Invariant Speech Recognition
Cory Stephenson, Jenelle Feather, Suchismita Padhy +4
Encouraged by the success of deep neural networks on a variety of visual tasks, much theoretical and experimental work has been aimed at understanding and interpreting how vision n…