2 citations · 2 across the 5 of their papers we have counts for
9 papers
Word Boundary Information Isn't Useful for Encoder Language Models
Edward Gow-Smith, Dylan Phelps, Harish Tayyar Madabushi +2
All existing transformer-based approaches to NLP using subword tokenisation algorithms encode whitespace (word boundary information) through the use of special space symbols (such…
Effective Cross-Task Transfer Learning for Explainable Natural Language Inference with T5
Irina Bigoulaeva, Rachneet Sachdeva, Harish Tayyar Madabushi +2
We compare sequential fine-tuning with a model for multi-task learning in the context where we are interested in boosting performance on two tasks, one of which depends on the othe…
SemEval-2022 Task 2: Multilingual Idiomaticity Detection and Sentence Embedding
Harish Tayyar Madabushi, Edward Gow-Smith, Marcos Garcia +3
This paper presents the shared task on Multilingual Idiomaticity Detection and Sentence Embedding, which consists of two subtasks: (a) a binary classification task aimed at identif…
Sample Efficient Approaches for Idiomaticity Detection
Dylan Phelps, Xuan-Rui Fan, Edward Gow-Smith +3
Deep neural models, in particular Transformer-based pre-trained language models, require a significant amount of data to train. This need for data tends to lead to problems when de…
AStitchInLanguageModels: Dataset and Methods for the Exploration of Idiomaticity in Pre-Trained Language Models
Harish Tayyar Madabushi, Edward Gow-Smith, Carolina Scarton +1
Despite their success in a variety of NLP tasks, pre-trained language models, due to their heavy reliance on compositionality, fail in effectively capturing the meanings of multiwo…
Investigating Language Impact in Bilingual Approaches for Computational Language Documentation
Marcely Zanon Boito, Aline Villavicencio, Laurent Besacier
For endangered languages, data collection campaigns have to accommodate the challenge that many of them are from oral tradition, and producing transcriptions is costly. Therefore,…