2 citations · 2 across the 5 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.CL2026
Building Legal Reward Models for Grounding and Abstention
Rilton Franzone, Valentin Noël, Puyu Wang +2
Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is in…
cs.AI2026
When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation
Griffin Farrow, Lily Sijia Li, Jack Johnson +4
Hallucinations can undermine clinician trust in LLMs, making it important that evaluation methods capture clinically relevant errors. Rubric-based evaluation has become the leading…