activity
20202025
most citedThe GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

52 citations · 56 across the 4 of their papers we have counts for

collaborators

6 papers

cs.AI20251 cited

Causes in neuron diagrams, and testing causal reasoning in Large Language Models. A glimpse of the future of philosophy?

Louis Vervoort, Vitaly Nikolaev

We propose a test for abstract causal reasoning in AI, based on scholarship in the philosophy of causation, in particular on the neuron diagrams popularized by D. Lewis. We illustr…

cs.CL20223 cited

TaTa: A Multilingual Table-to-Text Dataset for African Languages

Sebastian Gehrmann, Sebastian Ruder, Vitaly Nikolaev +4

Existing data-to-text generation datasets are mostly limited to English. To address this lack of data, we create Table-to-Text in African languages (TaTa), the first large multilin…

cs.CL2021

Planning with Learned Entity Prompts for Abstractive Summarization

Shashi Narayan, Yao Zhao, Joshua Maynez +3

We introduce a simple but flexible mechanism to learn an intermediate plan to ground the generation of abstractive summaries. Specifically, we prepend (or prompt) target summaries…

cs.CL202152 cited

The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal +53

We introduce GEM, a living benchmark for natural language Generation (NLG), its Evaluation, and Metrics. Measuring progress in NLG relies on a constantly evolving ecosystem of auto…

eess.IV2020

Convolution Neural Networks for Semantic Segmentation: Application to Small Datasets of Biomedical Images

Vitaly Nikolaev

This thesis studies how the segmentation results, produced by convolutional neural networks (CNN), is different from each other when applied to small biomedical datasets. We use di…

cs.CL2020

TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages

Jonathan H. Clark, Eunsol Choi, Michael Collins +4

Confidently making progress on multilingual modeling requires challenging, trustworthy evaluations. We present TyDi QA---a question answering dataset covering 11 typologically dive…