5 papers
APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain
Marie Escribe, Tharindu Ranasinghe, Amal Haddad Haddad +2
Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, especially in specialised and hi…
MUNIChus: Multilingual News Image Captioning Benchmark
Yuji Chen, Alistair Plum, Hansi Hettiarachchi +4
The goal of news image captioning is to generate captions by integrating news article content with corresponding images, highlighting the relationship between textual context and v…
AHaSIS: Shared Task on Sentiment Analysis for Arabic Dialects
Maram Alharbi, Salmane Chafik, Saad Ezzini +3
The hospitality industry in the Arab world increasingly relies on customer feedback to shape services, driving the need for advanced Arabic sentiment analysis tools. To address thi…
CODE-ACCORD: A Corpus of building regulatory data for rule generation towards automatic compliance checking
Hansi Hettiarachchi, Amna Dridi, Mohamed Medhat Gaber +11
Automatic Compliance Checking (ACC) within the Architecture, Engineering, and Construction (AEC) sector necessitates automating the interpretation of building regulations to achiev…
Overview of the First Workshop on Language Models for Low-Resource Languages (LoResLM 2025)
Hansi Hettiarachchi, Tharindu Ranasinghe, Paul Rayson +5
The first Workshop on Language Models for Low-Resource Languages (LoResLM 2025) was held in conjunction with the 31st International Conference on Computational Linguistics (COLING…