works on

From the 1 of 13 linked papers with an AI index.

collaborators

13 papers

cs.LG2026

Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance

Chao Peng, Zhiheng Lyu, Peijie Dong +2

The paper proposes a benchmark metric called the horizon residual to compare long-horizon task success against predictions from short-stage baselines, highlighting how performance…

cs.CL2026

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen +35

We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model lea…

cs.AI2026

CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection

Qihua Pan, Zhenheng Tang, Peijie Dong +4

Evaluating Large Language Models (LLMs) incurs prohibitive computational overhead during continuous development processes. While coreset selection accelerates evaluation, existing…

cs.AI2026

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning

Kai Liu, Peijie Dong, Xinchen Xie +5

The rapid progress of reasoning and agentic large language models (LLMs) has increased the demand for long-context inference, but self-attention (SA) scales quadratically with cont…

cs.CL2026

VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing

Xiaoyan Su, Peijie Dong, Zhenheng Tang +8

Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic tasks essential for profess…

cs.CE2026

Position: LLM Inference Should Be Evaluated as Energy-to-Token Production

Xiang Liu, Shimiao Yuan, Zhenheng Tang +5

LLM inference is still evaluated mainly as a model or software problem: accuracy, latency, throughput, and hardware utilization. This is incomplete. At deployment scale, the releva…