2 papers
cs.AI2026
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
Sunishchal Dev, Andrew Sloan, Joshua Kavner +2
We present the Judge Reliability Harness, an open source library for constructing validation suites that test the reliability of LLM judges. As LLM based scoring is widely deployed…
cs.LG2018
Applied Federated Learning: Improving Google Keyboard Query Suggestions
Timothy Yang, Galen Andrew, Hubert Eichner +5
Federated learning is a distributed form of machine learning where both the training data and model training are decentralized. In this paper, we use federated learning in a commer…