45 citations · 202 across the 17 of their papers we have counts for
6 papers · 1 filter
Which transformer architecture fits my data? A vocabulary bottleneck in self-attention
Noam Wies, Yoav Levine, Daniel Jannai +1
After their successful debut in natural language processing, Transformer architectures are now becoming the de-facto standard in many domains. An obstacle for their deployment over…
The Depth-to-Width Interplay in Self-Attention
Yoav Levine, Noam Wies, Or Sharir +2
Self-attention architectures, which are rapidly pushing the frontier in natural language processing, demonstrate a surprising depth-inefficient behavior: previous works indicate th…
On the Ethics of Building AI in a Responsible Manner
Shai Shalev-Shwartz, Shaked Shammah, Amnon Shashua
The AI-alignment problem arises when there is a discrepancy between the goals that a human designer specifies to an AI learner and a potential catastrophic outcome that does not re…
On the Sample Complexity of End-to-end Training vs. Semantic Abstraction Training
Shai Shalev-Shwartz, Amnon Shashua
We compare the end-to-end training approach to a modular approach in which a system is decomposed into semantically meaningful components. We focus on the sample complexity aspect,…
Tightening Fractional Covering Upper Bounds on the Partition Function for High-Order Region Graphs
Tamir Hazan, Jian Peng, Amnon Shashua
In this paper we present a new approach for tightening upper bounds on the partition function. Our upper bounds are based on fractional covering bounds on the entropy function, and…
Introduction to Machine Learning: Class Notes 67577
Amnon Shashua
Introduction to Machine learning covering Statistical Inference (Bayes, EM, ML/MaxEnt duality), algebraic and spectral methods (PCA, LDA, CCA, Clustering), and PAC learning (the Fo…