Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
League: Leaderboard Generation on Demand
Jian Wu, Jiayu Zhang, Dongyuan Li +5
This paper introduces Leaderboard Auto Generation (LAG), a novel and well-organized framework for automatic generation of leaderboards on a given research topic in rapidly evolving…
cs.CL2025
ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning
Shulin Huang, Linyi Yang, Yan Song +9
Evaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challeng…
cs.CL2024
Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark
Jian Wu, Linyi Yang, Zhen Wang +2
While Large Language Models (LLMs) excel in question-answering (QA) tasks, their real reasoning abilities on multiple evidence retrieval and integration on Multi-hop QA tasks remai…