4 papers · 1 filter
Code-Switched Language Identification is Harder Than You Think
Laurie Burchell, Alexandra Birch, Robert P. Thompson +1
Code switching (CS) is a very common phenomenon in written and spoken communication but one that is handled poorly by many natural language processing applications. Looking to the…
An Open Dataset and Model for Language Identification
Laurie Burchell, Alexandra Birch, Nikolay Bogoychev +1
Language identification (LID) is a fundamental step in many natural language processing pipelines. However, current LID systems are far from perfect, particularly on lower-resource…
The University of Edinburgh's Submission to the WMT22 Code-Mixing Shared Task (MixMT)
Faheem Kirefu, Vivek Iyer, Pinzhen Chen +1
The University of Edinburgh participated in the WMT22 shared task on code-mixed translation. This consists of two subtasks: i) generating code-mixed Hindi/English (Hinglish) text g…
Querent Intent in Multi-Sentence Questions
Laurie Burchell, Jie Chi, Tom Hosking +2
Multi-sentence questions (MSQs) are sequences of questions connected by relations which, unlike sequences of standalone questions, need to be answered as a unit. Following Rhetoric…