activity
20162024
most citedSemEval-2022 Task 2: Multilingual Idiomaticity Detection and Sentence Embedding

2 citations · 2 across the 5 of their papers we have counts for

collaborators

9 papers

cs.CL2024

Word Boundary Information Isn't Useful for Encoder Language Models

Edward Gow-Smith, Dylan Phelps, Harish Tayyar Madabushi +2

All existing transformer-based approaches to NLP using subword tokenisation algorithms encode whitespace (word boundary information) through the use of special space symbols (such…

cs.CL2022

Effective Cross-Task Transfer Learning for Explainable Natural Language Inference with T5

Irina Bigoulaeva, Rachneet Sachdeva, Harish Tayyar Madabushi +2

We compare sequential fine-tuning with a model for multi-task learning in the context where we are interested in boosting performance on two tasks, one of which depends on the othe…

cs.CL20222 cited

SemEval-2022 Task 2: Multilingual Idiomaticity Detection and Sentence Embedding

Harish Tayyar Madabushi, Edward Gow-Smith, Marcos Garcia +3

This paper presents the shared task on Multilingual Idiomaticity Detection and Sentence Embedding, which consists of two subtasks: (a) a binary classification task aimed at identif…

cs.CL2022

Sample Efficient Approaches for Idiomaticity Detection

Dylan Phelps, Xuan-Rui Fan, Edward Gow-Smith +3

Deep neural models, in particular Transformer-based pre-trained language models, require a significant amount of data to train. This need for data tends to lead to problems when de…

cs.CL2021

AStitchInLanguageModels: Dataset and Methods for the Exploration of Idiomaticity in Pre-Trained Language Models

Harish Tayyar Madabushi, Edward Gow-Smith, Carolina Scarton +1

Despite their success in a variety of NLP tasks, pre-trained language models, due to their heavy reliance on compositionality, fail in effectively capturing the meanings of multiwo…

cs.CL2020

Investigating Language Impact in Bilingual Approaches for Computational Language Documentation

Marcely Zanon Boito, Aline Villavicencio, Laurent Besacier

For endangered languages, data collection campaigns have to accommodate the challenge that many of them are from oral tradition, and producing transcriptions is costly. Therefore,…