1 paper · 1 filter
Surya Narayanan Hari, Matt Thomson
The introduction of the transformer architecture and the self-attention mechanism has led to an explosive production of language models trained on specific downstream tasks and dat…