Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Evaluating Large Language Models with Psychometrics
Yuan Li, Yue Huang, Hongyi Wang +4
Large Language Models (LLMs) have demonstrated exceptional capabilities in solving various tasks, progressively evolving into general-purpose assistants. The increasing integration…
cs.CL2024
FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free Guarantees
Fan Nie, Xiaotian Hou, Shuhang Lin +3
The propensity of Large Language Models (LLMs) to generate hallucinations and non-factual content undermines their reliability in high-stakes domains, where rigorous control over T…
cs.CL2024
TrustLLM: Trustworthiness in Large Language Models
Yue Huang, Lichao Sun, Haoran Wang +67
Large language models (LLMs), exemplified by ChatGPT, have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs prese…