activity
20192026
most citedTest-time Fourier Style Calibration for Domain Generalization

4 citations · 13 across the 22 of their papers we have counts for

collaborators
Showing cs.CLShow all

19 papers · 1 filter

cs.CL2026

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

Mingqi Gao, Anthony Sicilia, Weiyan Shi

Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inferenc…

cs.CL2026

Change My View? The Dynamics of Persuasion and Polarization in Online Discourse

David Freeborn, Malihe Alikani, Anthony Sicilia

Philosophical accounts of persuasion often assume that shared evidence and rational argumentation should lead to a convergence of views between peers, yet everyday discourse often…

cs.CL2025

Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation

Mert İnan, Anthony Sicilia, Alex Xie +3

Establishing shared goals is a fundamental step in human-AI communication. However, ambiguities can lead to outputs that seem correct but fail to reflect the speaker's intent. In t…

cs.CL20252 cited

Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity

Jiayi Zhang, Simon Yu, Derek Chong +4

Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collapse. Unlike prior work that attributes this effect to algorithmic limitations, we id…

cs.CL2025

Measuring How (Not Just Whether) VLMs Build Common Ground

Saki Imai, Mert İnan, Anthony Sicilia +1

Large vision language models (VLMs) increasingly claim reasoning skills, yet current benchmarks evaluate them in single-turn or question answering settings. However, grounding is a…

cs.CL2025

SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation

Saki Imai, Mert İnan, Anthony Sicilia +1

Evaluating sign language generation is often done through back-translation, where generated signs are first recognized back to text and then compared to a reference using text-base…