activity
20182021
most citedBOSS: Bayesian Optimization over String Spaces

20 citations · 24 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2021

Understanding who uses Reddit: Profiling individuals with a self-reported bipolar disorder diagnosis

Glorianna Jagfeld, Fiona Lobban, Paul Rayson +1

Recently, research on mental health conditions using public online data, including Reddit, has surged in NLP and health research but has not reported user characteristics, which ar…

cs.CL2021

MasakhaNER: Named Entity Recognition for African Languages

David Ifeoluwa Adelani, Jade Abbott, Graham Neubig +58

We take a step towards addressing the under-representation of the African continent in NLP research by creating the first large publicly available high-quality dataset for named en…

cs.CL2020

The National Corpus of Contemporary Welsh: Project Report | Y Corpws Cenedlaethol Cymraeg Cyfoes: Adroddiad y Prosiect

Dawn Knight, Steve Morris, Tess Fitzpatrick +3

This report provides an overview of the CorCenCC project and the online corpus resource that was developed as a result of work on the project. The report lays out the theoretical u…

cs.CL2020

Igbo-English Machine Translation: An Evaluation Benchmark

Ignatius Ezeani, Paul Rayson, Ikechukwu Onyenwe +2

Although researchers and practitioners are pushing the boundaries and enhancing the capacities of NLP tools and methods, works on African languages are lagging. A lot of focus on w…

cs.CL2019

In Search of Meaning: Lessons, Resources and Next Steps for Computational Analysis of Financial Discourse

Mahmoud El-Haj, Paul Rayson, Martin Walker +2

We critically assess mainstream accounting and finance research applying methods from computational linguistics (CL) to study financial discourse. We also review common themes and…

cs.CL2018

Using J-K fold Cross Validation to Reduce Variance When Tuning NLP Models

Henry B. Moss, David S. Leslie, Paul Rayson

K-fold cross validation (CV) is a popular method for estimating the true performance of machine learning models, allowing model selection and parameter tuning. However, the very pr…