activity
20242026
collaborators

9 papers

cs.CL2026

IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation

Bosi Wen, Yilin Niu, Cunxiang Wang +5

Instruction-following is a foundational capability of large language models (LLMs), with its improvement hinging on scalable and accurate feedback from judge models. However, the r…

cs.CL2025

Deep Research: A Systematic Survey

Zhengliang Shi, Yiqun Chen, Haitao Li +23

Large language models (LLMs) have rapidly evolved from text generators into powerful problem solvers. Yet, many open tasks demand critical thinking, multi-source, and verifiable ou…

cs.CL2025

Deep Literature Survey Automation with an Iterative Workflow

Hongbo Zhang, Han Cui, Yidong Wang +6

Automatic literature survey generation has attracted increasing attention, yet most existing systems follow a one-shot paradigm, where a large set of papers is retrieved at once an…

cs.AI2025

TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them

Yidong Wang, Yunze Song, Tingyuan Zhu +11

The adoption of Large Language Models (LLMs) as automated evaluators (LLM-as-a-judge) has revealed critical inconsistencies in current evaluation frameworks. We identify two fundam…

cs.CL2025

NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens

Cunxiang Wang, Ruoxi Ning, Boqi Pan +8

Recent advancements in Large Language Models (LLMs) have pushed the boundaries of natural language processing, especially in long-context understanding. However, the evaluation of…

cs.CL2025

LongRAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall

Zehan Qi, Rongwu Xu, Zhijiang Guo +3

Retrieval-augmented generation (RAG) is a promising approach to address the limitations of fixed knowledge in large language models (LLMs). However, current benchmarks for evaluati…