activity
20242026
collaborators

5 papers

cs.CL2026

R2-Write: Reflection and Revision for Open-Ended Writing with Deep Reasoning

Wanlong Liu, Bo Zhang, Chenliang Li +4

While deep reasoning with long chain-of-thought has dramatically improved large language models in verifiable domains like mathematics, its effectiveness for open-ended tasks such…

cs.AI2025

AgentBench: Evaluating LLMs as Agents

Xiao Liu, Hao Yu, Hanchen Zhang +19

The potential of Large Language Model (LLM) as agents has been widely acknowledged recently. Thus, there is an urgent need to quantitatively \textit{evaluate LLMs as agents} on cha…

cs.CL2024

AlignBench: Benchmarking Chinese Alignment of Large Language Models

Xiao Liu, Xuanyu Lei, Shengyuan Wang +15

Alignment has become a critical step for instruction-tuned Large Language Models (LLMs) to become helpful assistants. However, the effective evaluation of alignment for emerging Ch…

cs.CL2024

CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Pei Ke, Bosi Wen, Zhuoer Feng +9

Since the natural language processing (NLP) community started to make large language models (LLMs) act as a critic to evaluate the quality of generated texts, most of the existing…

cs.CL2024

SafetyBench: Evaluating the Safety of Large Language Models

Zhexin Zhang, Leqi Lei, Lindong Wu +7

With the rapid development of Large Language Models (LLMs), increasing attention has been paid to their safety concerns. Consequently, evaluating the safety of LLMs has become an e…