5 papers
Bootstrapping Embeddings for Low Resource Languages
Merve Basoz, Andrew Horne, Mattia Opper
Embedding models are crucial to modern NLP. However, the creation of the most effective models relies on carefully constructed supervised finetuning data. For high resource languag…
Mechanisms of Symbol Processing for In-Context Learning in Transformer Networks
Paul Smolensky, Roland Fernandez, Zhenghao Herbert Zhou +3
Large Language Models (LLMs) have demonstrated impressive abilities in symbol processing through in-context learning (ICL). This success flies in the face of decades of critiques a…
TRA: Better Length Generalisation with Threshold Relative Attention
Mattia Opper, Roland Fernandez, Paul Smolensky +1
Transformers struggle with length generalisation, displaying poor performance even on basic tasks. We test whether these limitations can be explained through two key failures of th…
Banyan: Improved Representation Learning with Explicit Structure
Mattia Opper, N. Siddharth
We present Banyan, a model that efficiently learns semantic representations by leveraging explicit hierarchical structure. While transformers excel at scale, they struggle in low-r…
Compositional Generalization Across Distributional Shifts with Sparse Tree Operations
Paul Soulos, Henry Conklin, Mattia Opper +3
Neural networks continue to struggle with compositional generalization, and this issue is exacerbated by a lack of massive pre-training. One successful approach for developing neur…