activity
20172026
most citedThe GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

52 citations · 214 across the 19 of their papers we have counts for

collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

Avijit Ghosh, Anka Reuel, Jenny Chim +45

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…

cs.AI2025

In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI

Shayne Longpre, Kevin Klyman, Ruth E. Appel +31

The widespread deployment of general-purpose AI (GPAI) systems introduces significant new risks. Yet the infrastructure, practices, and norms for reporting flaws in GPAI systems re…

cs.AI202311 cited

The Troubling Emergence of Hallucination in Large Language Models -- An Extensive Definition, Quantification, and Prescriptive Remediations

Vipula Rawte, Swagata Chakraborty, Agnibh Pathak +5

The recent advancements in Large Language Models (LLMs) have garnered widespread acclaim for their remarkable emerging capabilities. However, the issue of hallucination has paralle…

cs.AI2019

Why Build an Assistant in Minecraft?

Arthur Szlam, Jonathan Gray, Kavya Srinet +11

In this document we describe a rationale for a research program aimed at building an open "assistant" in the game Minecraft, in order to make progress on the problems of natural la…

cs.AI201918 cited

CraftAssist: A Framework for Dialogue-enabled Interactive Agents

Jonathan Gray, Kavya Srinet, Yacine Jernite +6

This paper describes an implementation of a bot assistant in Minecraft, and the tools and platform allowing players to interact with the bot and to record those interactions. The p…