1 citations · 1 across the 11 of their papers we have counts for
14 papers · 1 filter
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation
Katelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi +5
Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-…
ReLay: Personalized LLM-Generated Plain-Language Summaries for Better Understanding, but at What Cost?
Joey Chan, Yikun Han, Jingyuan Chen +8
Plain Language Summaries (PLS) aim to make research accessible to lay readers, but they are typically written in a one-size-fits-all style that ignores differences in readers' info…
CQA-Eval: Designing Reliable Evaluations of Multi-paragraph Clinical QA under Resource Constraints
Federica Bologna, Tiffany Pan, Matthew Wilkens +2
Evaluating multi-paragraph clinical question answering (QA) systems is resource-intensive and challenging: accurate judgments require medical expertise and achieving consistent hum…
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
Jena D. Hwang, Varsha Kishore, Amanpreet Singh +9
Recent advances have made long-form report-generating systems widely available. This has prompted evaluation frameworks that use LLM-as-judge protocols and claim verification, alon…
Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification
Zongwan Cao, Bingbing Wen, Lucy Lu Wang
Real-world visual question answering (VQA) is often context-dependent: an image-question pair may be under-specified, such that the correct answer depends on external information t…
ROBoto2: An Interactive System and Dataset for LLM-assisted Clinical Trial Risk of Bias Assessment
Anthony Hevia, Sanjana Chintalapati, Veronica Ka Wai Lai +4
We present ROBOTO2, an open-source, web-based platform for large language model (LLM)-assisted risk of bias (ROB) assessment of clinical trials. ROBOTO2 streamlines the traditional…