collaborators

6 papers

cs.CV2025

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

David Ma, Huaqing Yuan, Xingjian Wang +16

Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hour…

cs.CL2025

SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

P Team, Xinrun Du, Yifan Yao +94

Large language models (LLMs) have demonstrated remarkable proficiency in mainstream academic disciplines such as mathematics, physics, and computer science. However, human knowledg…

cs.CL2025

Multi-Agent Collaboration for Multilingual Code Instruction Tuning

Jian Yang, Wei Zhang, Jiaxi Yang +9

Recent advancement in code understanding and generation demonstrates that code LLMs fine-tuned on a high-quality instruction dataset can gain powerful capabilities to address wide-…

cs.CL2025

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Shanghaoran Quan, Jiaxi Yang, Bowen Yu +14

With the increasing code reasoning capabilities of existing large language models (LLMs) and breakthroughs in reasoning models like OpenAI o1 and o3, there is a growing need to dev…

cs.CL2025

Qwen2.5 Technical Report

Qwen, :, An Yang +41

In this report, we introduce Qwen2.5, a comprehensive series of large language models (LLMs) designed to meet diverse needs. Compared to previous iterations, Qwen 2.5 has been sign…

cs.CL2024

Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Bofei Gao, Feifan Song, Zhe Yang +17

Recent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH ar…