Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
UQ: Assessing Language Models on Unsolved Questions
Fan Nie, Ken Ziyu Liu, Zihao Wang +11
Benchmarks shape progress in AI research. A useful benchmark should be both difficult and realistic: questions should challenge frontier models while also reflecting real-world usa…
cs.CL2024
FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free Guarantees
Fan Nie, Xiaotian Hou, Shuhang Lin +3
The propensity of Large Language Models (LLMs) to generate hallucinations and non-factual content undermines their reliability in high-stakes domains, where rigorous control over T…