activity
20212024
most citedAll That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text

16 citations · 22 across the 6 of their papers we have counts for

collaborators

6 papers

cs.HC2024★ 1 cited

How Performance Pressure Influences AI-Assisted Decision Making

Nikita Haduong, Noah A. Smith

Many domains now employ AI-based decision-making aids, and although the potential for AI systems to assist with decision making is much discussed, human-AI collaboration often unde…

cs.CL2024★ 3 cited

Risks and NLP Design: A Case Study on Procedural Document QA

Nikita Haduong, Alice Gao, Noah A. Smith

As NLP systems are increasingly deployed at scale, concerns about their potential negative impacts have attracted the attention of the research community, yet discussions of risk h…

cs.HC2024

CPS-TaskForge: Generating Collaborative Problem Solving Environments for Diverse Communication Tasks

Nikita Haduong, Irene Wang, Bo-Ru Lu +2

Teams can outperform individuals; could adding AI teammates further bolster performance of teams solving problems collaboratively? Collaborative problem solving (CPS) research comm…

cs.CL2024

Efficient Encoder-Decoder Transformer Decoding for Decomposable Tasks

Bo-Ru Lu, Nikita Haduong, Chien-Yu Lin +3

Transformer-based NLP models are powerful but have high computational costs that limit deployment. Finetuned encoder-decoder models are popular in specialized domains and can outpe…

cs.CL2023★ 2 cited

Does Collaborative Human-LM Dialogue Generation Help Information Extraction from Human Dialogues?

Bo-Ru Lu, Nikita Haduong, Chia-Hsuan Lee +7

The capabilities of pretrained language models have opened opportunities to explore new application areas, but applications involving human-human interaction are limited by the fac…

cs.CL2021★ 16 cited

All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text

Elizabeth Clark, Tal August, Sofia Serrano +3

Human evaluations are typically considered the gold standard in natural language generation, but as models' fluency improves, how well can evaluators detect and judge machine-gener…