3 citations · 6 across the 4 of their papers we have counts for
4 papers
Pre-Trained Language Models Represent Some Geographic Populations Better Than Others
Jonathan Dunn, Benjamin Adams, Harish Tayyar Madabushi
This paper measures the skew in how well two families of LLMs represent diverse geographic populations. A spatial probing task is used with geo-referenced corpora to measure the de…
Abstraction not Memory: BERT and the English Article System
Harish Tayyar Madabushi, Dagmar Divjak, Petar Milin
Article prediction is a task that has long defied accurate linguistic description. As such, this task is ideally suited to evaluate models on their ability to emulate native-speake…
SemEval-2022 Task 2: Multilingual Idiomaticity Detection and Sentence Embedding
Harish Tayyar Madabushi, Edward Gow-Smith, Marcos Garcia +3
This paper presents the shared task on Multilingual Idiomaticity Detection and Sentence Embedding, which consists of two subtasks: (a) a binary classification task aimed at identif…
Improving Tokenisation by Alternative Treatment of Spaces
Edward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton +1
Tokenisation is the first step in almost all NLP tasks, and state-of-the-art transformer-based language models all use subword tokenisation algorithms to process input text. Existi…