From the 4 of 17 linked papers with an AI index.
14 papers · 1 filter
NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems
Jiayu Liu, Rui Wang, Qing Zong +9
Accurately assessing model confidence is essential for deploying large language models (LLMs) in mission-critical factual domains. While retrieval-augmented generation (RAG) is wid…
SurveyLens: A Discipline-Aware Benchmark for Automatic Survey Generation
Beichen Guo, Zhiyuan Wen, Jia Gu +6
Automatic Survey Generation (ASG) aims to produce comprehensive literature surveys by retrieving, organizing, and synthesizing academic papers. Despite rapid progress in specialize…
OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset
Wenbin Hu, Huihao Jing, Haochen Shi +3
Ensuring the safety and compliance of large language models (LLMs) is of paramount importance. However, existing LLM safety datasets often rely on ad-hoc taxonomies for data genera…
The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning
Tianshi Zheng, Yixiang Chen, Chengxi Li +7
Chain-of-Thought (CoT) prompting has been widely recognized for its ability to enhance reasoning capabilities in large language models (LLMs). However, our study reveals a surprisi…
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
Qing Zong, Jiayu Liu, Tianshi Zheng +7
Accurate confidence calibration in Large Language Models (LLMs) is critical for safe use in high-stakes domains, where clear verbalized confidence enhances user trust. Traditional…
Safety Compliance: Rethinking LLM Safety Reasoning through the Lens of Compliance
Wenbin Hu, Huihao Jing, Haochen Shi +2
The proliferation of Large Language Models (LLMs) has demonstrated remarkable capabilities, elevating the critical importance of LLM safety. However, existing safety methods rely o…