activity
20212026
most citedWho Goes First? Influences of Human-AI Workflow on Decision Making in Clinical Imaging

7 citations · 7 across the 8 of their papers we have counts for

collaborators

8 papers

cs.AI2026

Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse

Martin Bertran, Riccardo Fogliato, Zhiwei Steven Wu

Empirical conclusions depend not only on data but on analytic decisions made throughout the research process. Many-analyst studies have quantified this dependence: independent team…

cs.CL2025

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Evangelia Spiliopoulou, Riccardo Fogliato, Hanna Burnsky +4

Large language models (LLMs) can serve as judges that offer rapid and reliable assessments of other LLM outputs. However, models may systematically assign overly favorable ratings…

cs.CL2025

Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles

Kimberly Le Truong, Riccardo Fogliato, Hoda Heidari +1

Current benchmarks for evaluating Large Language Models (LLMs) often do not exhibit enough writing style diversity, with many adhering primarily to standardized conventions. Such b…

cs.LG2024

Precise Model Benchmarking with Only a Few Observations

Riccardo Fogliato, Pratik Patil, Nil-Jana Akpinar +1

How can we precisely estimate a large language model's (LLM) accuracy on questions belonging to a specific topic within a larger question-answering dataset? The standard direct est…

cs.CV2024

A Framework for Efficient Model Evaluation through Stratification, Sampling, and Estimation

Riccardo Fogliato, Pratik Patil, Mathew Monfort +1

Model performance evaluation is a critical and expensive task in machine learning and computer vision. Without clear guidelines, practitioners often estimate model accuracy using a…

stat.ME2023

Confidence Intervals for Error Rates in 1:1 Matching Tasks: Critical Statistical Analysis and Recommendations

Riccardo Fogliato, Pratik Patil, Pietro Perona

Matching algorithms are commonly used to predict matches between items in a collection. For example, in 1:1 face verification, a matching algorithm predicts whether two face images…