24 papers
OpenCompass: A Universal Evaluation Platform for Large Language Models
Maosong Cao, Kai Chen, Haodong Duan +27
In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the…
Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
Zhonghang Yuan, Zhefan Wang, Fang Hu +7
Reinforcement learning with verifiable rewards (RLVR) has demonstrated promising potential to enhance the reasoning capabilities of large language models (LLMs) in domains such as…
Rectifying LLM Thought from Lens of Optimization
Junnan Liu, Hongwei Liu, Songyang Zhang +1
Recent advancements in large language models (LLMs) have been driven by their emergent reasoning capabilities, particularly through long chain-of-thought (CoT) prompting, which ena…
Benchmarking and Boosting Multilingual Capabilities of LVLMs via OCR-Centric Reinforcement Learning
Junyuan Gao, Jiahe Song, Jiang Wu +11
Evaluating the multilingual capabilities of Large Vision-Language Models (LVLMs) remains challenging because most benchmarks rely on non-parallel corpora, making it unclear whether…
How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
Zihan Ma, Dongsheng Zhu, Shudong Liu +6
Current safety evaluations for LLM-driven agents primarily focus on atomic harms, failing to address sophisticated threats where malicious intent is concealed or diluted within com…
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
Mo Li, Songyang Zhang, Taolin Zhang +3
The capability of large language models to handle long-context information is crucial across various real-world applications. Existing evaluation methods often rely either on real-…