activity
20242026
most citedSusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents

1 citations · 1 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

14 papers · 1 filter

cs.CL2026

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation

Katelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi +5

Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-…

cs.CL2026

ReLay: Personalized LLM-Generated Plain-Language Summaries for Better Understanding, but at What Cost?

Joey Chan, Yikun Han, Jingyuan Chen +8

Plain Language Summaries (PLS) aim to make research accessible to lay readers, but they are typically written in a one-size-fits-all style that ignores differences in readers' info…

cs.CL2026

CQA-Eval: Designing Reliable Evaluations of Multi-paragraph Clinical QA under Resource Constraints

Federica Bologna, Tiffany Pan, Matthew Wilkens +2

Evaluating multi-paragraph clinical question answering (QA) systems is resource-intensive and challenging: accurate judgments require medical expertise and achieving consistent hum…

cs.CL2026

Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks

Jena D. Hwang, Varsha Kishore, Amanpreet Singh +9

Recent advances have made long-form report-generating systems widely available. This has prompted evaluation frameworks that use LLM-as-judge protocols and claim verification, alon…

cs.CL2026

Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification

Zongwan Cao, Bingbing Wen, Lucy Lu Wang

Real-world visual question answering (VQA) is often context-dependent: an image-question pair may be under-specified, such that the correct answer depends on external information t…

cs.CL2025

ROBoto2: An Interactive System and Dataset for LLM-assisted Clinical Trial Risk of Bias Assessment

Anthony Hevia, Sanjana Chintalapati, Veronica Ka Wai Lai +4

We present ROBOTO2, an open-source, web-based platform for large language model (LLM)-assisted risk of bias (ROB) assessment of clinical trials. ROBOTO2 streamlines the traditional…