activity
20202024
most citedReplacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

15 citations · 15 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL202415 cited

Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Pat Verga, Sebastian Hofstatter, Sophia Althammer +6

As Large Language Models (LLMs) have become more advanced, they have outpaced our abilities to accurately evaluate their quality. Not only is finding data to adequately probe parti…

cs.CL2021

Two Approaches to Building Collaborative, Task-Oriented Dialog Agents through Self-Play

Arkady Arkhangorodsky, Scot Fang, Victoria Knight +3

Task-oriented dialog systems are often trained on human/human dialogs, such as collected from Wizard-of-Oz interfaces. However, human/human corpora are frequently too small for sup…

cs.CL2021

MeetDot: Videoconferencing with Live Translation Captions

Arkady Arkhangorodsky, Christopher Chu, Scot Fang +5

We present MeetDot, a videoconferencing system with live translation captions overlaid on screen. The system aims to facilitate conversation between people who speak different lang…

cs.CL2020

MEEP: An Open-Source Platform for Human-Human Dialog Collection and End-to-End Agent Training

Arkady Arkhangorodsky, Amittai Axelrod, Christopher Chu +6

We create a new task-oriented dialog platform (MEEP) where agents are given considerable freedom in terms of utterances and API calls, but are constrained to work within a push-but…

cs.RO2020

Zeus: A System Description of the Two-Time Winner of the Collegiate SAE AutoDrive Competition

Keenan Burnett, Jingxing Qian, Xintong Du +14

The SAE AutoDrive Challenge is a three-year collegiate competition to develop a self-driving car by 2020. The second year of the competition was held in June 2019 at MCity, a mock…