20 citations · 24 across the 5 of their papers we have counts for
7 papers · 1 filter
Understanding who uses Reddit: Profiling individuals with a self-reported bipolar disorder diagnosis
Glorianna Jagfeld, Fiona Lobban, Paul Rayson +1
Recently, research on mental health conditions using public online data, including Reddit, has surged in NLP and health research but has not reported user characteristics, which ar…
MasakhaNER: Named Entity Recognition for African Languages
David Ifeoluwa Adelani, Jade Abbott, Graham Neubig +58
We take a step towards addressing the under-representation of the African continent in NLP research by creating the first large publicly available high-quality dataset for named en…
The National Corpus of Contemporary Welsh: Project Report | Y Corpws Cenedlaethol Cymraeg Cyfoes: Adroddiad y Prosiect
Dawn Knight, Steve Morris, Tess Fitzpatrick +3
This report provides an overview of the CorCenCC project and the online corpus resource that was developed as a result of work on the project. The report lays out the theoretical u…
Igbo-English Machine Translation: An Evaluation Benchmark
Ignatius Ezeani, Paul Rayson, Ikechukwu Onyenwe +2
Although researchers and practitioners are pushing the boundaries and enhancing the capacities of NLP tools and methods, works on African languages are lagging. A lot of focus on w…
In Search of Meaning: Lessons, Resources and Next Steps for Computational Analysis of Financial Discourse
Mahmoud El-Haj, Paul Rayson, Martin Walker +2
We critically assess mainstream accounting and finance research applying methods from computational linguistics (CL) to study financial discourse. We also review common themes and…
Using J-K fold Cross Validation to Reduce Variance When Tuning NLP Models
Henry B. Moss, David S. Leslie, Paul Rayson
K-fold cross validation (CV) is a popular method for estimating the true performance of machine learning models, allowing model selection and parameter tuning. However, the very pr…