most citedOn the geometry of generalization and memorization in deep neural networks

13 citations · 32 across the 4 of their papers we have counts for

collaborators

5 papers

cs.LG202113 cited

On the geometry of generalization and memorization in deep neural networks

Cory Stephenson, Suchismita Padhy, Abhinav Ganesh +3

Understanding how large neural networks avoid memorizing training data is key to explaining their high generalization performance. To examine the structure of when and where memori…

cs.CL20211 cited

Syntactic Perturbations Reveal Representational Correlates of Hierarchical Phrase Structure in Pretrained Language Models

Matteo Alleman, Jonathan Mamou, Miguel A Del Rio +3

While vector-based language representations from pretrained language models have set a new standard for many NLP tasks, there is not yet a complete accounting of their inner workin…

cs.LG2021

1-bit LAMB: Communication Efficient Large-Scale Large-Batch Training with LAMB's Convergence Speed

Conglong Li, Ammar Ahmad Awan, Hanlin Tang +2

To train large models (like BERT and GPT-3) on hundreds of GPUs, communication has become a major bottleneck, especially on commodity systems with limited-bandwidth TCP network. On…

cs.CL202010 cited

Emergence of Separable Manifolds in Deep Language Representations

Jonathan Mamou, Hang Le, Miguel Del Rio +4

Deep neural networks (DNNs) have shown much empirical success in solving perceptual tasks across various cognitive modalities. While they are only loosely inspired by the biologica…

cs.LG20208 cited

Untangling in Invariant Speech Recognition

Cory Stephenson, Jenelle Feather, Suchismita Padhy +4

Encouraged by the success of deep neural networks on a variety of visual tasks, much theoretical and experimental work has been aimed at understanding and interpreting how vision n…