3 papers
cs.AI2026
TruthTensor: Evaluating LLMs through Human Imitation on Prediction Market under Drift and Holistic Reasoning
Shirin Shahabi, Spencer Graham, Haruna Isah
Evaluating language models and AI agents remains fundamentally challenging because static benchmarks fail to capture real-world uncertainty, distribution shift, and the gap between…
cs.CR2025
JSTprove: Pioneering Verifiable AI for a Trustless Future
Jonathan Gold, Tristan Freiberg, Haruna Isah +1
The integration of machine learning (ML) systems into critical industries such as healthcare, finance, and cybersecurity has transformed decision-making processes, but it also brin…
cs.AI2025
DSperse: A Framework for Targeted Verification in Zero-Knowledge Machine Learning
Dan Ivanov, Tristan Freiberg, Shirin Shahabi +2
DSperse is a modular framework for distributed machine learning inference with strategic cryptographic verification. Operating within the emerging paradigm of distributed zero-know…