3 papers
cs.CL2026
SAFE: An LLM-as-Verifier Framework for Evidence-Grounded Multi-Hop Reasoning
Daeyong Kwon, Soyoung Yoon, Seung-won Hwang
Multi-hop QA benchmarks often reward Large Language Models (LLMs) for spurious correctness, where models reach correct answers through invalid intermediate reasoning. We propose SA…
cs.CL2025
GRADE: Generating multi-hop QA and fine-gRAined Difficulty matrix for RAG Evaluation
Jeongsoo Lee, Daeyong Kwon, Kyohoon Jin
Retrieval-Augmented Generation (RAG) systems are widely adopted in knowledge-intensive NLP tasks, but current evaluations often overlook the structural complexity and multi-step re…
cs.IR2025
MHTS: Multi-Hop Tree Structure Framework for Generating Difficulty-Controllable QA Datasets for RAG Evaluation
Jeongsoo Lee, Daeyong Kwon, Kyohoon Jin +3
Existing RAG benchmarks often overlook query difficulty, leading to inflated performance on simpler questions and unreliable evaluations. A robust benchmark dataset must satisfy th…