most citedLongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

2 citations · 3 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2026

On the Paradoxical Interference between Instruction-Following and Task Solving

Yunjia Qi, Hao Peng, Xintong Shi +5

Instruction following aims to align Large Language Models (LLMs) with human intent by specifying explicit constraints on how tasks should be performed. However, we reveal a counter…

cs.CL20251 cited

StoryWriter: A Multi-Agent Framework for Long Story Generation

Haotian Xia, Hao Peng, Yunjia Qi +4

Long story generation remains a challenge for existing large language models (LLMs), primarily due to two main factors: (1) discourse coherence, which requires plot consistency, lo…

cs.CL2025

VerIF: Verification Engineering for Reinforcement Learning in Instruction Following

Hao Peng, Yunjia Qi, Xiaozhi Wang +3

Reinforcement learning with verifiable rewards (RLVR) has become a key technique for enhancing large language models (LLMs), with verification engineering playing a central role. H…

cs.AI2025

AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Yunjia Qi, Hao Peng, Xiaozhi Wang +5

Large Language Models (LLMs) have demonstrated advanced capabilities in real-world agentic applications. Growing research efforts aim to develop LLM-based agents to address practic…

cs.CL2025

MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark

Shengkun Ma, Hao Peng, Lei Hou +1

Machine Reading Comprehension (MRC) is an essential task in evaluating natural language understanding. Existing MRC datasets primarily assess specific aspects of reading comprehens…

cs.CL20252 cited

LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Yushi Bai, Shangqing Tu, Jiajie Zhang +9

This paper introduces LongBench v2, a benchmark designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world…