21 papers
OpenCompass: A Universal Evaluation Platform for Large Language Models
Maosong Cao, Kai Chen, Haodong Duan +27
In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the…
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
Zhonghang Yuan, Zhefan Wang, Fang Hu +7
Reinforcement learning with verifiable rewards (RLVR) has demonstrated promising potential to enhance the reasoning capabilities of large language models (LLMs) in domains such as…
Rectifying LLM Thought from Lens of Optimization
Junnan Liu, Hongwei Liu, Songyang Zhang +1
Recent advancements in large language models (LLMs) have been driven by their emergent reasoning capabilities, particularly through long chain-of-thought (CoT) prompting, which ena…
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
Junyuan Gao, Jiahe Song, Jiang Wu +11
While Large Vision-Language Models (LVLMs) demonstrate promising multilingual capabilities, their evaluation is currently hindered by two critical limitations: (1) the use of non-p…
How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
Zihan Ma, Dongsheng Zhu, Shudong Liu +6
Current safety evaluations for LLM-driven agents primarily focus on atomic harms, failing to address sophisticated threats where malicious intent is concealed or diluted within com…
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
Mo Li, Songyang Zhang, Taolin Zhang +3
The capability of large language models to handle long-context information is crucial across various real-world applications. Existing evaluation methods often rely either on real-…