3 papers
cs.CL2026
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
cs.LG2026
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
Tavor Z. Baharav, Spyros Dragazis, Aldo Pacchiano
We study sequential testing for a binary disease outcome when risk follows an unknown logistic model. At each round, the decision maker may either pay for a test revealing the true…
cs.DS2025
Estimating Hitting Times Locally At Scale
Themistoklis Haris, Fabian Spaeh, Spyros Dragazis +1
Hitting times provide a fundamental measure of distance in random processes, quantifying the expected number of steps for a random walk starting at node to reach node . They…