9 citations · 9 across the 3 of their papers we have counts for
6 papers · 1 filter
RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
Dongyu Ru, Lin Qiu, Xiangkun Hu +15
Despite Retrieval-Augmented Generation (RAG) showing promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to th…
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
Zhen Huang, Zengzhi Wang, Shijie Xia +25
The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showc…
Prompt Chaining or Stepwise Prompt? Refinement in Text Summarization
Shichao Sun, Ruifeng Yuan, Ziqiang Cao +2
Large language models (LLMs) have demonstrated the capacity to improve summary quality by mirroring a human-like iterative process of critique and refinement starting from the init…
Dissecting Human and LLM Preferences
Junlong Li, Fan Zhou, Shichao Sun +3
As a relative quality comparison of model responses, human and Large Language Model (LLM) preferences serve as common alignment goals in model fine-tuning and criteria in evaluatio…
The Critique of Critique
Shichao Sun, Junlong Li, Weizhe Yuan +3
Critique, as a natural language description for assessing the quality of model-generated content, has played a vital role in the training, evaluation, and refinement of LLMs. Howev…
Generative Judge for Evaluating Alignment
Junlong Li, Shichao Sun, Weizhe Yuan +3
The rapid development of Large Language Models (LLMs) has substantially expanded the range of tasks they can address. In the field of Natural Language Processing (NLP), researchers…