Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality
Wen Luo, Guangyue Peng, Liang Wang +7
Large Reasoning Models achieve strong performance on complex tasks but remain prone to hallucinations, particularly in long-form generation where errors compound across reasoning s…
cs.CL2024
HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation
Wen Luo, Tianshu Shen, Wei Li +4
Large Language Models (LLMs) have significantly advanced the field of Natural Language Processing (NLP), achieving remarkable performance across diverse tasks and enabling widespre…