activity
20242026
most citedQwen2.5 Technical Report

107 citations · 167 across the 12 of their papers we have counts for

collaborators

14 papers

cs.SE2026

SWE-PolyVision: Benchmarking Cross-Image Abductive Reasoning for Repository-Level Software Engineering

Jiajun Wu, Leixin Sun, Zihan Tan +9

Current multimodal software-engineering benchmarks expose images as additional context, but do not test whether an agent can integrate evidence distributed across images into a ver…

cs.AI2026

SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

Jiajun Wu, Leixin Sun, Zihan Tan +9

Large language model based coding agents have made substantial progress on repository-level software engineering tasks. Existing repository benchmarks, however, usually start from…

cs.CV2025

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

David Ma, Huaqing Yuan, Xingjian Wang +16

Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hour…

cs.CL2025★ 4 cited

SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

P Team, Xinrun Du, Yifan Yao +94

Large language models (LLMs) have demonstrated remarkable proficiency in mainstream academic disciplines such as mathematics, physics, and computer science. However, human knowledg…

cs.CL2025

Multi-Agent Collaboration for Multilingual Code Instruction Tuning

Jian Yang, Wei Zhang, Jiaxi Yang +9

Recent advancement in code understanding and generation demonstrates that code LLMs fine-tuned on a high-quality instruction dataset can gain powerful capabilities to address wide-…

cs.CL2025★ 4 cited

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Shanghaoran Quan, Jiaxi Yang, Bowen Yu +14

With the increasing code reasoning capabilities of existing large language models (LLMs) and breakthroughs in reasoning models like OpenAI o1 and o3, there is a growing need to dev…