4 papers
NumLeak: Public Numeric Benchmarks as Latent Labels in Foundation Models
Anany Kotawala
Public numeric benchmarks appear in pretraining, so an evaluation that conditions on a date may be measuring memorized recall rather than out-of-sample skill. We introduce NumLeak,…
Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents
Anany Kotawala
Multi-component LLM agents assemble probabilistic claims from components that each see only part of a joint problem; the composition can violate basic probability axioms even when…
Resolution Diagnostics for Paired LLM Evaluation
Anany Kotawala
Across two public LLM leaderboards, many displayed pairwise rankings do not meet a conventional paired-test resolution target under the actual paired evaluation design: 11 of 40 Op…
GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models
Vartan Shadarevian, Kia Ghods, Alex Kenich +1
Large language models (LLMs) are increasingly deployed as economic agents in marketplaces, auctions, and bidding settings. Anticipating their behavior in any specific deployment is…