6 papers
RepEval: Effective Text Evaluation with LLM Representation
Shuqian Sheng, Yi Xu, Tianhang Zhang +7
The era of Large Language Models (LLMs) raises new demands for automatic evaluation metrics, which should be adaptable to various application scenarios while maintaining low cost a…
SH2: Self-Highlighted Hesitation Helps You Decode More Truthfully
Jushi Kai, Tianhang Zhang, Hai Hu +1
Large language models (LLMs) demonstrate great performance in text generation. However, LLMs are still suffering from hallucinations. In this work, we propose an inference-time met…
ECon: On the Detection and Resolution of Evidence Conflicts
Cheng Jiayang, Chunkit Chan, Qianqian Zhuang +7
The rise of large language models (LLMs) has significantly influenced the quality of information in decision-making systems, leading to the prevalence of AI-generated content and c…
RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
Dongyu Ru, Lin Qiu, Xiangkun Hu +15
Despite Retrieval-Augmented Generation (RAG) showing promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to th…
RefChecker: Reference-based Fine-grained Hallucination Checker and Benchmark for Large Language Models
Xiangkun Hu, Dongyu Ru, Lin Qiu +7
Large Language Models (LLMs) have shown impressive capabilities but also a concerning tendency to hallucinate. This paper presents RefChecker, a framework that introduces claim-tri…
GeoGalactica: A Scientific Large Language Model in Geoscience
Zhouhan Lin, Cheng Deng, Le Zhou +18
Large language models (LLMs) have achieved huge success for their general knowledge and ability to solve a wide spectrum of tasks in natural language processing (NLP). Due to their…