2 papers
cs.AI2026
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
Sunishchal Dev, Andrew Sloan, Joshua Kavner +2
We present the Judge Reliability Harness, an open source library for constructing validation suites that test the reliability of LLM judges. As LLM based scoring is widely deployed…
cs.GT2025
Bridging Theory and Perception in Fair Division: A Study on Comparative and Fair Share Notions
Hadi Hosseini, Joshua Kavner, Samarth Khanna +2
The allocation of resources among multiple agents is a fundamental problem in both economics and computer science. In these settings, fairness plays a crucial role in ensuring soci…