4 papers · 1 filter
Learning interpretable positional encodings in transformers depends on initialization
Takuya Ito, Luca Cocchi, Tim Klinger +3
In transformers, the positional encoding (PE) provides essential information that distinguishes the position and order amongst tokens in a sequence. Most prior investigations of PE…
Transformers Learn Faster with Semantic Focus
Parikshit Ram, Kenneth L. Clarkson, Tim Klinger +2
Various forms of sparse attention have been explored to mitigate the quadratic computational and memory cost of the attention mechanism in transformers. We study sparse transformer…
Neural Reasoning Networks: Efficient Interpretable Neural Networks With Automatic Textual Explanations
Stephen Carrow, Kyle Harper Erwin, Olga Vilenskaia +5
Recent advances in machine learning have led to a surge in adoption of neural networks for various tasks, but lack of interpretability remains an issue for many others in which an…
What makes Models Compositional? A Theoretical View: With Supplement
Parikshit Ram, Tim Klinger, Alexander G. Gray
Compositionality is thought to be a key component of language, and various compositional benchmarks have been developed to empirically probe the compositional generalization of exi…