3 papers
cs.CL2023
Embedding structure matters: Comparing methods to adapt multilingual vocabularies to new languages
C. M. Downey, Terra Blevins, Nora Goldfine +1
Pre-trained multilingual language models underpin a large portion of modern NLP tools outside of English. A strong baseline for specializing these models for specific languages is…
cs.CL2023
Evaluating Transformer's Ability to Learn Mildly Context-Sensitive Languages
Shunjie Wang, Shane Steinert-Threlkeld
Despite the fact that Transformers perform well in NLP tasks, recent studies suggest that self-attention is theoretically limited in learning even some regular and context-free lan…
cs.LG2023
The Weighted Möbius Score: A Unified Framework for Feature Attribution
Yifan Jiang, Shane Steinert-Threlkeld
Feature attribution aims to explain the reasoning behind a black-box model's prediction by identifying the impact of each feature on the prediction. Recent work has extended featur…