6 papers
ASSERT: A Measurement Pipeline for GenAI Audits
Riccardo Fogliato, Abhinav Palia, Xiawei Wang +11
Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate t…
SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents
Rui Melo, Riccardo Fogliato, Sean Zhou +2
Large language models (LLMs) are increasingly deployed in automated code-review systems, where their approvals can determine which code is merged into shared repositories. However,…
Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality
Xiaoyuan Zhu, Kimberly Le Truong, Riccardo Fogliato +8
As LLMs are deployed in high-stakes settings, users must judge the correctness of individual responses, often relying on model-generated justifications such as reasoning chains or…
Stronger Neyman Regret Guarantees for Adaptive Experimental Design
Georgy Noarov, Riccardo Fogliato, Martin Bertran +1
We study the design of adaptive, sequential experiments for unbiased average treatment effect (ATE) estimation in the design-based potential outcomes setting. Our goal is to develo…
Improving LLM Group Fairness on Tabular Data via In-Context Learning
Valeriia Cherepanova, Chia-Jung Lee, Nil-Jana Akpinar +4
Large language models (LLMs) have been shown to be effective on tabular prediction tasks in the low-data regime, leveraging their internal knowledge and ability to learn from instr…
Multicalibration for Confidence Scoring in LLMs
Gianluca Detommaso, Martin Bertran, Riccardo Fogliato +1
This paper proposes the use of "multicalibration" to yield interpretable and reliable confidence scores for outputs generated by large language models (LLMs). Multicalibration asks…