activity
20242026
collaborators

8 papers

cs.SE2026

Transformer-Assisted LLM-Based Source Code Summarisation: to Enable More Secure Software Development

Jesse Phillips, Tracy Hall, Paul Rayson +1

Neural Source Code Summarisation (NSCS) aims to generate natural language summaries of source code to improve developers' and maintainers' understanding of code. Source code summar…

cs.CL2026

GaelEval: Benchmarking LLM Performance for Scottish Gaelic

Peter Devine, William Lamb, Beatrice Alex +5

Multilingual large language models (LLMs) often exhibit emergent 'shadow' capabilities in languages without official support, yet their performance on these languages remains uneve…

cs.CL2026

Structured Prompting for Arabic Essay Proficiency: A Trait-Centric Evaluation Approach

Salim Al Mandhari, Hieu Pham Dinh, Mo El-Haj +1

This paper presents a novel prompt engineering framework for trait specific Automatic Essay Scoring (AES) in Arabic, leveraging large language models (LLMs) under zero-shot and few…

cs.CL2026

Creating a Hybrid Rule and Neural Network Based Semantic Tagger using Silver Standard Data: the PyMUSAS framework for Multilingual Semantic Annotation

Andrew Moore, Paul Rayson, Dawn Archer +10

Word Sense Disambiguation (WSD) has been widely evaluated using the semantic frameworks of WordNet, BabelNet, and the Oxford Dictionary of English. However, for the UCREL Semantic…

cs.CL2026

FreeTxt-Vi: A Benchmarked Vietnamese-English Toolkit for Segmentation, Sentiment, and Summarisation

Hung Nguyen Huy, Mo El-Haj, Dawn Knight +1

FreeTxt-Vi is a free and open source web based toolkit for creating and analysing bilingual Vietnamese English text collections. Positioned at the intersection of corpus linguistic…

cs.CL2025

HealthcareNLP: where are we and what is next?

Lifeng Han, Paul Rayson, Suzan Verberne +2

This proposed tutorial focuses on Healthcare Domain Applications of NLP, what we have achieved around HealthcareNLP, and the challenges that lie ahead for the future. Existing revi…