15 papers
Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation
Yongkang Yang, Zhezheng Hao, Hong Zhang +8
On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-distillation. Two recent research…
Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance
Chao Peng, Zhiheng Lyu, Peijie Dong +2
The paper proposes a benchmark metric called the horizon residual to compare long-horizon task success against predictions from short-stage baselines, highlighting how performance…
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen +35
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model lea…
LEPO: Latent Reasoning Policy Optimization for Large Language Models
Yuyan Zhou, Jiarui Yu, Hande Dong +4
Recently, latent reasoning has been introduced into large language models (LLMs) to leverage rich information within a continuous space. However, without stochastic sampling, these…
Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems
Zhezheng Hao, Tianfu Wang, Huanshuo Dong +7
LLM-based multi-agent systems (MAS) have emerged as an effective paradigm for complex and long-horizon tasks. However, in real-world tasks, MAS often exhibit various failures durin…
Echo: Learning from Experience Data via User-Driven Refinement
Hande Dong, Xiaoyun Liang, Jiarui Yu +15
Static "human data" faces inherent limitations: it is expensive to scale and bounded by the knowledge of its creators. Continuous learning from "experience data" - interactions bet…