107 citations · 167 across the 12 of their papers we have counts for
14 papers
SWE-PolyVision: Benchmarking Cross-Image Abductive Reasoning for Repository-Level Software Engineering
Jiajun Wu, Leixin Sun, Zihan Tan +9
Current multimodal software-engineering benchmarks expose images as additional context, but do not test whether an agent can integrate evidence distributed across images into a ver…
SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories
Jiajun Wu, Leixin Sun, Zihan Tan +9
Large language model based coding agents have made substantial progress on repository-level software engineering tasks. Existing repository benchmarks, however, usually start from…
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
David Ma, Huaqing Yuan, Xingjian Wang +16
Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hour…
SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
P Team, Xinrun Du, Yifan Yao +94
Large language models (LLMs) have demonstrated remarkable proficiency in mainstream academic disciplines such as mathematics, physics, and computer science. However, human knowledg…
Multi-Agent Collaboration for Multilingual Code Instruction Tuning
Jian Yang, Wei Zhang, Jiaxi Yang +9
Recent advancement in code understanding and generation demonstrates that code LLMs fine-tuned on a high-quality instruction dataset can gain powerful capabilities to address wide-…
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Shanghaoran Quan, Jiaxi Yang, Bowen Yu +14
With the increasing code reasoning capabilities of existing large language models (LLMs) and breakthroughs in reasoning models like OpenAI o1 and o3, there is a growing need to dev…