activity
20202025
most citedAILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

4 citations · 5 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2025

Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning

Jinlong Liu, Mohammed Bahja, Venelin Kovatchev +1

Evaluating and optimising authorial style in long-form story generation remains challenging because style is often assessed with ad hoc prompting and is frequently conflated with o…

cs.CL2024

Benchmark Transparency: Measuring the Impact of Data on Evaluation

Venelin Kovatchev, Matthew Lease

In this paper we present an exploratory research on quantifying the impact that data distribution has on the performance and evaluation of NLP models. We propose an automated frame…

cs.CL20221 cited

InferES : A Natural Language Inference Corpus for Spanish Featuring Negation-Based Contrastive and Adversarial Examples

Venelin Kovatchev, Mariona Taulé

In this paper, we present InferES - an original corpus for Natural Language Inference (NLI) in European Spanish. We propose, implement, and analyze a variety of corpus-creating str…

cs.CL2022

ProtoTEx: Explaining Model Decisions with Prototype Tensors

Anubrata Das, Chitrank Gupta, Venelin Kovatchev +2

We present ProtoTEx, a novel white-box NLP classification architecture based on prototype networks. ProtoTEx faithfully explains model decisions based on prototype tensors that enc…

cs.CL2021

Can vectors read minds better than experts? Comparing data augmentation strategies for the automated scoring of children's mindreading ability

Venelin Kovatchev, Phillip Smith, Mark Lee +1

In this paper we implement and compare 7 different data augmentation strategies for the task of automatic scoring of children's ability to understand others' thoughts, feelings, an…

cs.CL2020

"What is on your mind?" Automated Scoring of Mindreading in Childhood and Early Adolescence

Venelin Kovatchev, Phillip Smith, Mark Lee +3

In this paper we present the first work on the automated scoring of mindreading ability in middle childhood and early adolescence. We create MIND-CA, a new corpus of 11,311 questio…