4 citations · 13 across the 22 of their papers we have counts for
19 papers · 1 filter
Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation
Mingqi Gao, Anthony Sicilia, Weiyan Shi
Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inferenc…
Change My View? The Dynamics of Persuasion and Polarization in Online Discourse
David Freeborn, Malihe Alikani, Anthony Sicilia
Philosophical accounts of persuasion often assume that shared evidence and rational argumentation should lead to a convergence of views between peers, yet everyday discourse often…
Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation
Mert İnan, Anthony Sicilia, Alex Xie +3
Establishing shared goals is a fundamental step in human-AI communication. However, ambiguities can lead to outputs that seem correct but fail to reflect the speaker's intent. In t…
Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
Jiayi Zhang, Simon Yu, Derek Chong +4
Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collapse. Unlike prior work that attributes this effect to algorithmic limitations, we id…
Measuring How (Not Just Whether) VLMs Build Common Ground
Saki Imai, Mert İnan, Anthony Sicilia +1
Large vision language models (VLMs) increasingly claim reasoning skills, yet current benchmarks evaluate them in single-turn or question answering settings. However, grounding is a…
SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation
Saki Imai, Mert İnan, Anthony Sicilia +1
Evaluating sign language generation is often done through back-translation, where generated signs are first recognized back to text and then compared to a reference using text-base…