4 papers
Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias
Justin D. Norman, Michael U. Rivera, D. Alex Hughes
LLM-as-a-Judge has become the dominant evaluation paradigm for language models, but judge validation in practice relies on exact-match agreement, a metric that does not correct for…
Does Head Pose Correction Improve Biometric Facial Recognition?
Justin Norman, Hany Farid
Biometric facial recognition models often demonstrate significant decreases in accuracy when processing real-world images, often characterized by poor quality, non-frontal subject…
The Case for Repeatable, Open, and Expert-Grounded Hallucination Benchmarks in Large Language Models
Justin D. Norman, Michael U. Rivera, D. Alex Hughes
Plausible, but inaccurate, tokens in model-generated text are widely believed to be pervasive and problematic for the responsible adoption of language models. Despite this concern,…
Detecting Deepfake Talking Heads from Facial Biometric Anomalies
Justin D. Norman, Hany Farid
The combination of highly realistic voice cloning, along with visually compelling avatar, face-swap, or lip-sync deepfake video generation, makes it relatively easy to create a vid…