activity
20242026
collaborators

24 papers

cs.CL2026

OpenCompass: A Universal Evaluation Platform for Large Language Models

Maosong Cao, Kai Chen, Haodong Duan +27

In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the…

cs.CL2026

Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains

Zhonghang Yuan, Zhefan Wang, Fang Hu +7

Reinforcement learning with verifiable rewards (RLVR) has demonstrated promising potential to enhance the reasoning capabilities of large language models (LLMs) in domains such as…

cs.CL2026

Rectifying LLM Thought from Lens of Optimization

Junnan Liu, Hongwei Liu, Songyang Zhang +1

Recent advancements in large language models (LLMs) have been driven by their emergent reasoning capabilities, particularly through long chain-of-thought (CoT) prompting, which ena…

cs.CV2026

Benchmarking and Boosting Multilingual Capabilities of LVLMs via OCR-Centric Reinforcement Learning

Junyuan Gao, Jiahe Song, Jiang Wu +11

Evaluating the multilingual capabilities of Large Vision-Language Models (LVLMs) remains challenging because most benchmarks rely on non-parallel corpora, making it unclear whether…

cs.MA2025

How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity

Zihan Ma, Dongsheng Zhu, Shudong Liu +6

Current safety evaluations for LLM-driven agents primarily focus on atomic harms, failing to address sophisticated threats where malicious intent is concealed or diluted within com…

cs.CL2025

NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities

Mo Li, Songyang Zhang, Taolin Zhang +3

The capability of large language models to handle long-context information is crucial across various real-world applications. Existing evaluation methods often rely either on real-…