activity
20242026
collaborators

6 papers

cs.CL2026

ASSERT: A Measurement Pipeline for GenAI Audits

Riccardo Fogliato, Abhinav Palia, Xiawei Wang +11

Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate t…

cs.CR2026

SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents

Rui Melo, Riccardo Fogliato, Sean Zhou +2

Large language models (LLMs) are increasingly deployed in automated code-review systems, where their approvals can determine which code is merged into shared repositories. However,…

cs.HC2026

Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality

Xiaoyuan Zhu, Kimberly Le Truong, Riccardo Fogliato +8

As LLMs are deployed in high-stakes settings, users must judge the correctness of individual responses, often relying on model-generated justifications such as reasoning chains or…

stat.ME2025

Stronger Neyman Regret Guarantees for Adaptive Experimental Design

Georgy Noarov, Riccardo Fogliato, Martin Bertran +1

We study the design of adaptive, sequential experiments for unbiased average treatment effect (ATE) estimation in the design-based potential outcomes setting. Our goal is to develo…

cs.LG2024

Improving LLM Group Fairness on Tabular Data via In-Context Learning

Valeriia Cherepanova, Chia-Jung Lee, Nil-Jana Akpinar +4

Large language models (LLMs) have been shown to be effective on tabular prediction tasks in the low-data regime, leveraging their internal knowledge and ability to learn from instr…

stat.ML2024

Multicalibration for Confidence Scoring in LLMs

Gianluca Detommaso, Martin Bertran, Riccardo Fogliato +1

This paper proposes the use of "multicalibration" to yield interpretable and reliable confidence scores for outputs generated by large language models (LLMs). Multicalibration asks…