4 papers
CultureScope: A Dimensional Lens for Probing Cultural Understanding in LLMs
Jinghao Zhang, Sihang Jiang, Shiwei Guo +7
As large language models (LLMs) are increasingly deployed in diverse cultural environments, evaluating their cultural understanding capability has become essential for ensuring tru…
INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance
Shisong Chen, Qian Zhu, Wenyan Yang +15
Insurance, as a critical component of the global financial system, demands high standards of accuracy and reliability in AI applications. While existing benchmarks evaluate AI capa…
Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation
Jinyi Han, Tingyun Li, Shisong Chen +8
While large language models (LLMs) have demonstrated remarkable performance across diverse tasks, they fundamentally lack self-awareness and frequently exhibit overconfidence, assi…
Enhancing Confidence Expression in Large Language Models Through Learning from Past Experience
Haixia Han, Tingyun Li, Shisong Chen +5
Large Language Models (LLMs) have exhibited remarkable performance across various downstream tasks, but they may generate inaccurate or false information with a confident tone. One…