most citedQwen2.5 Technical Report

107 citations · 123 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2025

AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models

Qin Zhu, Fei Huang, Runyu Peng +6

While logical reasoning evaluation of Large Language Models (LLMs) has attracted significant attention, existing benchmarks predominantly rely on multiple-choice formats that are v…

cs.CL202512 cited

Qwen2.5-1M Technical Report

An Yang, Bowen Yu, Chengyuan Li +25

We introduce Qwen2.5-1M, a series of models that extend the context length to 1 million tokens. Compared to the previous 128K version, the Qwen2.5-1M series have significantly enha…

cs.CL2025

RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques

Zhengyang Tang, Ziniu Li, Zhenyang Xiao +8

Critiques are important for enhancing the performance of Large Language Models (LLMs), enabling both self-improvement and constructive feedback for others by identifying flaws and…

cs.CL20254 cited

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Shanghaoran Quan, Jiaxi Yang, Bowen Yu +14

With the increasing code reasoning capabilities of existing large language models (LLMs) and breakthroughs in reasoning models like OpenAI o1 and o3, there is a growing need to dev…

cs.CL2025107 cited

Qwen2.5 Technical Report

Qwen, :, An Yang +41

In this report, we introduce Qwen2.5, a comprehensive series of large language models (LLMs) designed to meet diverse needs. Compared to previous iterations, Qwen 2.5 has been sign…