activity
20242026
collaborators

5 papers

cs.CL2026

APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain

Marie Escribe, Tharindu Ranasinghe, Amal Haddad Haddad +2

Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, especially in specialised and hi…

cs.CL2026

MUNIChus: Multilingual News Image Captioning Benchmark

Yuji Chen, Alistair Plum, Hansi Hettiarachchi +4

The goal of news image captioning is to generate captions by integrating news article content with corresponding images, highlighting the relationship between textual context and v…

cs.CL2025

AHaSIS: Shared Task on Sentiment Analysis for Arabic Dialects

Maram Alharbi, Salmane Chafik, Saad Ezzini +3

The hospitality industry in the Arab world increasingly relies on customer feedback to shape services, driving the need for advanced Arabic sentiment analysis tools. To address thi…

cs.IR2025

CODE-ACCORD: A Corpus of building regulatory data for rule generation towards automatic compliance checking

Hansi Hettiarachchi, Amna Dridi, Mohamed Medhat Gaber +11

Automatic Compliance Checking (ACC) within the Architecture, Engineering, and Construction (AEC) sector necessitates automating the interpretation of building regulations to achiev…

cs.CL2024

Overview of the First Workshop on Language Models for Low-Resource Languages (LoResLM 2025)

Hansi Hettiarachchi, Tharindu Ranasinghe, Paul Rayson +5

The first Workshop on Language Models for Low-Resource Languages (LoResLM 2025) was held in conjunction with the 31st International Conference on Computational Linguistics (COLING…