6 papers · 1 filter
RefuteBench 2.0 -- Agentic Benchmark for Dynamic Evaluation of LLM Responses to Refutation Instruction
Jianhao Yan, Yun Luo, Yue Zhang
In the multi-turn interaction schema, large language models (LLMs) can leverage user feedback to enhance the quality and relevance of their responses. However, evaluating an LLM's…
Keys to Robust Edits: from Theoretical Insights to Practical Advances
Jianhao Yan, Futing Wang, Yun Luo +2
Large language models (LLMs) struggle with maintaining accurate knowledge due to conflicting/outdated parametric memories. While locate-and-edit methods address this, their relianc…
RefChecker: Reference-based Fine-grained Hallucination Checker and Benchmark for Large Language Models
Xiangkun Hu, Dongyu Ru, Lin Qiu +7
Large Language Models (LLMs) have shown impressive capabilities but also a concerning tendency to hallucinate. This paper presents RefChecker, a framework that introduces claim-tri…
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models
Jianhao Yan, Yun Luo, Yue Zhang
The application scope of large language models (LLMs) is increasingly expanding. In practical use, users might provide feedback based on the model's output, hoping for a responsive…
Enhancing Argument Structure Extraction with Efficient Leverage of Contextual Information
Yun Luo, Zhen Yang, Fandong Meng +3
Argument structure extraction (ASE) aims to identify the discourse structure of arguments within documents. Previous research has demonstrated that contextual information is crucia…
XAL: EXplainable Active Learning Makes Classifiers Better Low-resource Learners
Yun Luo, Zhen Yang, Fandong Meng +5
Active learning (AL), which aims to construct an effective training set by iteratively curating the most formative unlabeled data for annotation, has been widely used in low-resour…