8 papers
DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models
Yixin Bu, Runze Xia, Guanyun Zou +3
Accurate Uncertainty Quantification (UQ) is critical for reliable deployment of Large Language Models (LLMs), yet traditional probability-based metrics often fail to capture the mo…
Diversity-Oriented Fine-Tuning for Uncertainty-Based Hallucination Detection
Qiuyuan Li, Hongliang Dai, Piji Li
Existing hallucination detection methods are typically conducted at the inference stage, without making any modifications to the model itself. In this paper, we are interested in e…
Investigating Advanced Reasoning of Large Language Models via Black-Box Environment Interaction
Congchi Yin, Tianyi Wu, Yankai Shu +5
Existing tasks fall short in evaluating reasoning ability of Large Language Models (LLMs) in an interactive, unknown environment. This deficiency leads to the isolated assessment o…
CRISP: Compressing Redundancy in Chain-of-Thought via Intrinsic Saliency Pruning
Yangsong Lan, Hongliang Dai, Piji Li
Long Chain-of-Thought (CoT) reasoning is pivotal for the success of recent reasoning models but suffers from high computational overhead and latency. While prior works attempt to c…
Hallucination Mitigating for Medical Report Generation
Ruoqing Zhao, Runze Xia, Piji Li
In the realm of medical report generation (MRG), the integration of natural language processing has emerged as a vital tool to alleviate the workload of radiologists. Despite the i…
Concise and Sufficient Sub-Sentence Citations for Retrieval-Augmented Generation
Guo Chen, Qiuyuan Li, Qiuxian Li +3
In retrieval-augmented generation (RAG) question answering systems, generating citations for large language model (LLM) outputs enhances verifiability and helps users identify pote…